Skip to content

Load balancer Services

A Kubernetes Service of type=LoadBalancer provisions a managed load balancer in the cluster’s VPC. Once the annotations are valid, kubectl get svc shows an EXTERNAL-IP within 30 to 90 seconds. Each port in the Service’s spec.ports becomes one frontend (a listening port) on the load balancer and one backend forwarding to the cluster’s ready nodes at that port’s NodePort. Annotations on the Service control everything else: the plan, placement, protocol, TLS, timeouts, health checks, access and routing.

The load balancer bills hourly under the plan you name, the same as one created from the panel. See Load balancers for the underlying product.

Every example uses the default annotation prefix service.beta.kubernetes.io/managed-loadbalancer-. The prefix is operator-configurable: K8S_ANNOTATION_PREFIX on the management server must match ANNOTATION_PREFIX on the cluster’s cloud controller manager.

  • A running cluster. See Clusters.
  • The name of a load balancer plan your provider offers in the cluster’s region. Every Service needs it in the plan annotation.
  • kubectl access to the cluster. See Access and security.

plan names the load balancer plan (the sizing tier) to deploy. It is required on every Service of this type.

metadata:
annotations:
service.beta.kubernetes.io/managed-loadbalancer-plan: std-1g

When it is missing, the Service is rejected with a Kubernetes Event carrying the message annotation 'service.beta.kubernetes.io/managed-loadbalancer-plan' is required. Run kubectl describe svc <name> and read the event at the bottom. When the name does not exist or is not enabled, the event says unknown_lb_plan. When the plan is not offered in the cluster’s region, it says lb_plan_not_in_region. See Troubleshooting.

By default the load balancer gets a public IP. Four annotations change that:

  • internal: "true": a VPC-internal load balancer with no public IP. EXTERNAL-IP reports its VPC address. Needs a private subnet in the cluster’s VPC and an active NAT gateway on that subnet.
  • vpc-only: "true": the same request, second spelling.
  • public: "false": the same request, third spelling. The default is true.
  • public-ip: "203.0.113.10": deploy on one specific static IP reserved on your account instead of a pool address. See Static IPs.

Contradiction rule: vpc-only: "true" combined with an explicit public: "true" is rejected with an Event rather than guessed. internal or vpc-only combined with public-ip is rejected too; the two are mutually exclusive. A public-ip that is not a reserved static IP on your account is also rejected.

backend-protocol picks the working mode for every frontend:

  • tcp (default): L4 passthrough. Pods see the original bytes, including TLS if any.
  • http: L7. Enables TLS termination at the load balancer, HTTP health checks, and path and host routing.
  • https: L7 with re-encryption; the backends serve TLS themselves. Rare.

Unrecognised values fall back to tcp. port-{N}-backend-protocol overrides the choice for a single frontend port.

Two rules apply:

  • A port 443 frontend in http mode with no ssl-mode for that port is downgraded to tcp passthrough. This stops the load balancer from serving plain HTTP where clients expect TLS; your backend’s own TLS becomes the visible behavior.
  • UDP Service ports cannot pass through the load balancer and are skipped.

ssl-mode picks how frontends terminate TLS: none (default), letsencrypt, inline, or certificate.

  • letsencrypt requires ssl-domain. The domain’s DNS record must point at the load balancer’s public IP. The certificate is issued once the domain resolves, and renewal is automatic.
  • inline requires ssl-cert and ssl-key: the fullchain certificate and the private key, each base64-encoded PEM. Encode them with base64 -w0 fullchain.pem and base64 -w0 privkey.pem. ssl-domain is an optional label.
  • certificate requires ssl-certificate-id: one or more IDs of certificates that already exist on your account’s Certificates page, comma-separated. The first is the default certificate and the rest are served by hostname (SNI) in the order given. The private key never leaves the platform and nothing is copied into the Service.

The per-port variants port-{N}-ssl-mode, port-{N}-ssl-domain, port-{N}-ssl-cert, port-{N}-ssl-key and port-{N}-ssl-certificate-id override the global keys for one port, so a single Service can serve Let’s Encrypt on 443 and an inline certificate on 8443. Precedence everywhere on this page: the per-port key first, then the global key, then the default.

Port 80 exception: the global SSL keys apply to every port except port 80. Port 80 stays plaintext unless you set port-80-ssl-mode explicitly. This keeps http:// working and leaves port 80 free for the redirect below and for Let’s Encrypt renewals.

ssl-redirect: "true" adds an HTTP to HTTPS 301 redirect on HTTP-mode frontends whose port is not 443. It needs backend-protocol: http and a port 80 in spec.ports to redirect from. The redirect skips the Let’s Encrypt challenge path, so renewals keep working.

Reusing a certificate from the Certificates page

Section titled “Reusing a certificate from the Certificates page”
metadata:
annotations:
service.beta.kubernetes.io/managed-loadbalancer-plan: std-1g
service.beta.kubernetes.io/managed-loadbalancer-backend-protocol: http
service.beta.kubernetes.io/managed-loadbalancer-ssl-mode: certificate
service.beta.kubernetes.io/managed-loadbalancer-ssl-certificate-id: 7f0c1d2e-0000-4000-8000-000000000001,7f0c1d2e-0000-4000-8000-000000000002

Every ID must belong to your account and be usable: not expired, and not a failed or still-pending Let’s Encrypt request. A certificate that does not meet this is rejected the same way whether it is missing or belongs to someone else. Unlike the other modes, a bad certificate reference does reject the Service update with invalid_annotation_value, naming the annotation and the ID, and the load balancer keeps the certificates it already has.

Removing the annotation, or switching to another ssl-mode, detaches the certificates from the load balancer but never deletes them. A certificate that is attached to a load balancer cannot be deleted until you detach it. Applying the same annotations again changes nothing and does not reload the load balancer.

An invalid or incomplete TLS annotation in the other modes does not reject the Service. The load balancer still comes up, but the affected port gets no certificate. Check your annotations when a port serves plaintext or the wrong certificate.

idle-timeout sets the seconds an idle connection stays open before the load balancer closes it. It applies to every frontend the Service creates and raises the timeout on the backend side of those frontends too. port-{N}-idle-timeout overrides it for one port.

Defaults when unset: 50 seconds for HTTP and HTTPS frontends, 3600 seconds for TCP frontends. The valid range is 30-86400. An out-of-range value rejects the Service with a Kubernetes Event that names the key.

Raise the timeout for long-lived connections such as websockets, database sessions and SSH:

metadata:
annotations:
service.beta.kubernetes.io/managed-loadbalancer-idle-timeout: "3600"
service.beta.kubernetes.io/managed-loadbalancer-port-5432-idle-timeout: "86400"

Here a websocket application keeps idle connections open for one hour on every port, while PostgreSQL on port 5432 gets 24 hours.

Two more annotations control the backend side directly, independent of idle-timeout:

  • connect-timeout: how long the load balancer waits for a connection to a backend pod to establish, in seconds. Range 1-75. Defaults to 5 seconds when unset.
  • server-timeout: how long a backend pod has once connected to send its response, in seconds. Range 1-86400. When unset it is derived from idle-timeout (the pre-existing behavior); an explicit server-timeout always wins over that derivation.

Both accept the port-{N}- per-port form, same precedence as every other key on this page: per-port, then global, then the default.

metadata:
annotations:
service.beta.kubernetes.io/managed-loadbalancer-connect-timeout: "10"
service.beta.kubernetes.io/managed-loadbalancer-server-timeout: "120"
service.beta.kubernetes.io/managed-loadbalancer-port-8443-server-timeout: "600"

Here every backend gets 10 seconds to accept a connection and 2 minutes to respond, except port 8443 (a slow upload endpoint), which gets 10 minutes to respond.

If you are migrating an Ingress from nginx-ingress, its proxy-*-timeout annotations map onto these two HAProxy-backed settings:

nginx-ingress annotation This platform Notes
nginx.ingress.kubernetes.io/proxy-connect-timeout connect-timeout Same meaning: time to establish the backend connection.
nginx.ingress.kubernetes.io/proxy-read-timeout server-timeout HAProxy has no separate read timeout; server-timeout covers it.
nginx.ingress.kubernetes.io/proxy-send-timeout server-timeout Same field as proxy-read-timeout above - HAProxy has no separate send timeout either, so one setting covers both directions.
nginx.ingress.kubernetes.io/proxy-body-size (no equivalent) HAProxy applies no request-body size limit at all. Uploads that needed a raised proxy-body-size on nginx-ingress already work here with no setting to configure.

The load balancer actively probes backends and removes failing ones from rotation. Nine annotations tune the probes:

  • health-check-enabled: true by default. Set "false" to stop active probing.
  • health-check-protocol: tcp or http. Defaults to the frontend’s mode.
  • health-check-port: where to probe. Defaults to the traffic port (the NodePort). Also accepts the literal traffic-port.
  • health-check-path: the HTTP path to probe. Default /. Ignored for tcp checks.
  • health-check-interval: seconds between probes. Default 5.
  • health-check-timeout: per-probe timeout in seconds. Default 3.
  • health-check-healthy-threshold: consecutive successes before a backend counts as up. Default 2.
  • health-check-unhealthy-threshold: consecutive failures before a backend counts as down. Default 3.
  • health-check-expect: http only. A response match such as status 200 or string OK. Unset by default.

When the Service sets externalTrafficPolicy: Local, checks default to the dedicated health check node port Kubernetes allocates, so nodes without a local pod drop out of rotation automatically. An explicit health-check-port still wins.

Three more annotations add passive checks: with passive-check-enabled (default false), the load balancer also watches real traffic for errors, and a backend that reaches passive-check-error-limit observed errors (default 10) gets the passive-check-on-error action, which by default marks it down. The accepted actions are mark-down (default), fail-check, sudden-death and fastinter; unrecognised values fall back to mark-down.

source-ranges is a comma-separated list of CIDRs (or single IPs) allowed to reach the load balancer. IPv4 and IPv6 are both accepted:

service.beta.kubernetes.io/managed-loadbalancer-source-ranges: "10.0.0.0/8,203.0.113.7/32"

The platform enforces it with a managed security group on the load balancer instance, named k8s-lb-src- plus the first 8 characters of the load balancer ID. Do not edit that group; it is rebuilt from the annotation on every update. Removing the annotation deletes the group and restores unrestricted access. A partially invalid list applies the valid entries and skips the rest. A fully invalid list rejects the Service with an Event.

security-groups is a comma-separated list of security group names or UUIDs from your account; mixing both forms is fine. They are attached to the load balancer instance. Unknown entries are skipped. The list is re-applied on every Service update, so changes propagate without recreating the load balancer. See Security groups.

routing-rules is a JSON array. Each rule binds to one frontend port and matches incoming traffic by SNI, path or host:

[
{"port": 443, "match": "sni", "value": "shop.example.com", "ssl_domain": "shop.example.com"}
]
  • port: required. A port that exists in spec.ports.
  • match: required. sni, path or host.
  • value: required. The domain for sni and host, the URL prefix for path.
  • ssl_domain: sni rules only. A per-rule Let’s Encrypt domain. It is issued separately and served by hostname on the same port.
  • certificate_id: sni rules only. The ID of a certificate on your Certificates page, served for this hostname. It cannot be combined with ssl_domain, and it follows the same ownership and usability rules as ssl-certificate-id.

The global ssl-domain provides the default certificate, served when a client sends no SNI or an unknown hostname. Reconciliation: setting the annotation to [] deletes every rule on the load balancer. Removing the annotation entirely leaves existing rules alone. Rules left out of a new JSON are deleted on the ports the annotation touches.

metadata:
annotations:
service.beta.kubernetes.io/managed-loadbalancer-routing-rules: |
[
{"port": 443, "match": "sni", "value": "web.example.com", "ssl_domain": "web.example.com"},
{"port": 443, "match": "sni", "value": "shop.example.com", "ssl_domain": "shop.example.com"}
]

traffic-split spreads traffic across multiple child Services with relative weights, for blue/green and canary releases. The cluster’s cloud controller manager reads the annotation, resolves each child Service to its allocated NodePort, and the management server applies the weights on the load balancer. Two forms:

Form Annotation Applies to Match
Standalone traffic-split Every frontend port on the Service None: a catch-all evaluated after every routing-rules rule
Embedded routing-rules[].backends One routing rule The rule’s match and value

Both forms share the same child object:

[
{"service": "app-blue", "namespace": "production", "weight": 80},
{"service": "app-green", "namespace": "production", "weight": 20}
]
  • service: required. The name of the child Service in the same cluster.
  • namespace: optional. Defaults to the parent Service’s namespace.
  • weight: required. A relative integer 0-1000, not a percentage: [80, 20] and [400, 100] behave identically. weight: 0 keeps the child in the configuration but drains it, the standard canary rollback.

Blue/green cut-over, standalone form. Cut over by flipping the weights to [0, 100] in one kubectl apply:

apiVersion: v1
kind: Service
metadata:
name: web
annotations:
service.beta.kubernetes.io/managed-loadbalancer-plan: std-1g
service.beta.kubernetes.io/managed-loadbalancer-backend-protocol: http
service.beta.kubernetes.io/managed-loadbalancer-traffic-split: |
[
{"service": "web-blue", "weight": 100},
{"service": "web-green", "weight": 0}
]
spec:
type: LoadBalancer
selector:
app: web
ports:
- name: http
port: 80
targetPort: 8080
protocol: TCP

The split applies to every port in spec.ports.

Canary, embedded form. Add a backends array to the rule that should split:

service.beta.kubernetes.io/managed-loadbalancer-routing-rules: |
[
{
"port": 443,
"match": "host",
"value": "admin.example.com",
"backends": [
{"service": "admin-v1", "weight": 95},
{"service": "admin-v2", "weight": 5}
]
},
{
"port": 443,
"match": "host",
"value": "www.example.com"
}
]

Requests for admin.example.com split 95/5 between admin-v1 and admin-v2. Requests for www.example.com keep serving the parent Service’s own pods.

Rules:

  • Each child needs an allocated NodePort. ClusterIP Services do not get one; make each child a NodePort or LoadBalancer Service. A child without an allocated NodePort is skipped with a warning in the cloud controller manager’s log.
  • A child in another namespace than the parent Service is refused unless your provider enabled cross-namespace splits.
  • Removing the traffic-split annotation removes the split, and the parent Service’s own pods serve all traffic again. Setting it to [] does the same.
  • The parent Service’s own selector serves traffic only if the parent is listed among the children. Otherwise the parent acts purely as the load balancer anchor.

Splits need a cloud controller manager image that includes the traffic-split controller, and a management server that implements the split endpoint. On an older management server the annotation is parsed but never applied; see Common problems below. For the full JSON grammar of both forms, see the cloud controller manager annotation reference.

A web application: plain HTTP on port 80 with an automatic redirect, Let’s Encrypt on port 443, health checks on a real endpoint, a one-hour idle timeout, and restricted source access:

apiVersion: v1
kind: Service
metadata:
name: web
annotations:
service.beta.kubernetes.io/managed-loadbalancer-plan: std-1g
service.beta.kubernetes.io/managed-loadbalancer-backend-protocol: http
service.beta.kubernetes.io/managed-loadbalancer-ssl-mode: letsencrypt
service.beta.kubernetes.io/managed-loadbalancer-ssl-domain: web.example.com
service.beta.kubernetes.io/managed-loadbalancer-ssl-redirect: "true"
service.beta.kubernetes.io/managed-loadbalancer-health-check-path: /healthz
service.beta.kubernetes.io/managed-loadbalancer-idle-timeout: "3600"
service.beta.kubernetes.io/managed-loadbalancer-source-ranges: "10.0.0.0/8,203.0.113.7/32"
service.beta.kubernetes.io/managed-loadbalancer-security-groups: web-allow-443
spec:
type: LoadBalancer
selector:
app: web
ports:
- name: http
port: 80
targetPort: 8080
protocol: TCP
- name: https
port: 443
targetPort: 8080
protocol: TCP

This produces a public load balancer with:

  • A port 80 frontend that stays plaintext and 301-redirects to HTTPS, except on the Let’s Encrypt challenge path.
  • A port 443 frontend terminating TLS for web.example.com and forwarding plain HTTP to the pods.
  • Health checks on /healthz every 5 seconds; a backend that fails 3 checks in a row leaves rotation.
  • Idle connections closed after one hour instead of 50 seconds.
  • Access restricted to 10.0.0.0/8 and 203.0.113.7, plus the rules of the web-allow-443 security group.

Point the DNS record for web.example.com at the EXTERNAL-IP the Service receives. The certificate is issued once the domain resolves.

The management server reads the keys below. Prefix every key with the annotation prefix. Boolean keys accept true, 1, yes, on and false, 0, no, off, case-insensitive. A port-{N}-<key> key overrides its global counterpart for frontend port {N}, where {N} is a spec.ports[].port value. Precedence: the per-port key, then the global key, then the default listed here.

Key Type Default Applies to Notes
plan string none, required Whole Service An enabled load balancer plan offered in the cluster’s region. Missing or invalid rejects the Service with an Event.
internal boolean false Placement VPC-internal, no public IP. Needs a private subnet in the cluster VPC. Mutually exclusive with public-ip.
vpc-only boolean false Placement Same as internal. With an explicit public: "true" it rejects the Service with an Event.
public boolean true Placement "false" deploys VPC-internal.
public-ip string pool address Placement Must be a static IP reserved on your account, otherwise rejected.
backend-protocol tcp, http, https tcp All frontends http enables TLS termination, HTTP checks, path and host routing. Unrecognised values fall back to tcp. Port 443 in http mode without SSL downgrades to tcp.
ssl-mode none, letsencrypt, inline, certificate none All frontends except port 80 Port 80 stays plaintext unless port-80-ssl-mode is set.
ssl-domain FQDN none TLS Required for letsencrypt; must resolve to the load balancer’s public IP. Optional label for inline.
ssl-cert base64 PEM none TLS Fullchain certificate. Required for inline.
ssl-key base64 PEM none TLS Private key. Required for inline.
ssl-certificate-id comma-separated UUIDs none TLS Account certificates for certificate mode. First is the default, the rest are added for SNI in order. Required for certificate.
ssl-redirect boolean false HTTP frontends other than 443 301 to HTTPS. Skips the Let’s Encrypt challenge path.
idle-timeout integer seconds 50 http/https, 3600 tcp All frontends Range 30-86400; out-of-range rejects the Service with an Event. Raises the backend-side timeout too.
connect-timeout integer seconds 5 Backends Time to establish the backend connection. Range 1-75; out-of-range rejects the Service with an Event.
server-timeout integer seconds derived from idle-timeout Backends Time a backend has to respond once connected. Range 1-86400; out-of-range rejects the Service with an Event. Covers both proxy-read-timeout and proxy-send-timeout from nginx-ingress. An explicit value wins over the idle-timeout derivation.
proxy-protocol v1, v2, boolean false Backends Sends the PROXY protocol header to backends (v2 binary, v1 text). Backends must parse the header or every request fails.
health-check-enabled boolean true Backends "false" stops active probing.
health-check-protocol tcp, http frontend’s mode Backends http enables health-check-path and health-check-expect.
health-check-port integer or traffic-port traffic port Backends Under externalTrafficPolicy: Local the default becomes the Kubernetes health check node port.
health-check-path string / Backends http only. Ignored for tcp checks.
health-check-interval integer seconds 5 Backends Probe frequency.
health-check-timeout integer seconds 3 Backends Per-probe timeout.
health-check-healthy-threshold integer 2 Backends Consecutive successes before a backend counts as up.
health-check-unhealthy-threshold integer 3 Backends Consecutive failures before a backend counts as down.
health-check-expect string unset Backends http only. Example: status 200, string OK.
passive-check-enabled boolean false Backends Watches real traffic for errors. Works with or without active probing.
passive-check-error-limit integer 10 Backends Observed errors tolerated before the passive-check-on-error action. Clamped to a minimum of 1.
passive-check-on-error mark-down, fail-check, sudden-death, fastinter mark-down Backends Action taken at the error limit. Unrecognised values fall back to mark-down.
source-ranges CSV of CIDRs unrestricted Access IPv4 and IPv6. Enforced by a managed security group; removing the annotation deletes the group.
security-groups CSV of names or UUIDs none Access Groups from your account. Unknown entries are skipped. Re-applied on every update.
routing-rules JSON array none Routing sni, path and host rules. [] deletes every rule; removing the key leaves rules alone. A rule’s optional backends array splits that rule’s traffic; see Traffic split.
traffic-split JSON array none All frontends (catch-all rule) Read by the cluster’s cloud controller manager and applied by the management server; needs a cloud controller manager image that includes the traffic-split controller.
port-{N}-backend-protocol as backend-protocol global key, then tcp Frontend port N Per-port mode override.
port-{N}-ssl-mode as ssl-mode global key, then none Frontend port N Per-port TLS mode.
port-{N}-ssl-domain as ssl-domain global key, then none Frontend port N Per-port domain.
port-{N}-ssl-cert as ssl-cert global key, then none Frontend port N Per-port inline certificate.
port-{N}-ssl-key as ssl-key global key, then none Frontend port N Per-port inline key.
port-{N}-ssl-certificate-id as ssl-certificate-id global key, then none Frontend port N Per-port account certificates.
port-{N}-idle-timeout as idle-timeout global key, then the mode default Frontend port N Same range and rejection behavior as idle-timeout.
port-{N}-connect-timeout as connect-timeout global key, then 5 Frontend port N Per-port connect timeout.
port-{N}-server-timeout as server-timeout global key, then the idle-timeout derivation Frontend port N Per-port response timeout.

The complete routing-rules and traffic-split grammar is documented in the cloud controller manager annotation reference.

  • EXTERNAL-IP stays <pending>. Run kubectl describe svc <name> and read the event at the bottom; its message names the exact cause. A missing or unknown plan is the most common one. See Troubleshooting.
  • Requests through the load balancer return 503. Every backend failed its health check and left rotation. The default HTTP check probes /, and an application that returns 404 there fails the check even while healthy. Point health-check-path at a real endpoint such as /healthz.
  • TLS on port 443 fails. ssl-mode is absent or incomplete, so the load balancer passes TCP through instead of terminating TLS and clients get no certificate from it. Set ssl-mode with its companion keys and re-apply the Service.
  • Websocket, database or SSH connections drop while idle. The default idle timeouts apply: 50 seconds for HTTP and HTTPS frontends, 3600 for TCP. Raise idle-timeout, or port-{N}-idle-timeout for the affected port.
  • A slow endpoint gets a 502/504 before it finishes responding. The backend ran past server-timeout (or its idle-timeout-derived default). Raise server-timeout, or port-{N}-server-timeout for the affected port.
  • Large uploads fail or truncate. This is not a proxy-body-size equivalent - there isn’t one to configure. HAProxy applies no request-body size limit, so check server-timeout instead: a slow upload that outruns it looks like a truncated request.
  • The split never applies. Run kubectl -n kube-system logs deploy/cloud-controller-manager | grep traffic-split. Repeated reconcile errors carrying the management server’s 501 not_implemented answer mean the management server is older than the traffic-split feature; the controller retries with backoff, so the split applies on its own once the management server is updated.