Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Kubernetes as a service works when a tenant can declare what they need in one small API object, and a controller keeps the cluster and any external systems moving toward that declaration. The custom resource is the service contract. The controller is the reconciliation engine that creates and maintains the Kubernetes objects and external infrastructure behind it.

Most of the hard decisions follow from that split: which extension mechanism exposes the contract, how the controller reports progress, how tenants are isolated, who owns networking, and which framework you use to write the loop. This guide stays architectural. It does not assume a cloud provider, tenancy model, or service-level objective, and it flags where a choice depends on those.

How the pieces fit together

A platform of this kind has four parts. Keeping them separate in your design makes ownership and failure handling easier to reason about.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Layer What it does Typical owner
Service contract A custom resource with a spec for desired state and a status for observed state Platform team designs it; tenants create instances
Controller A control loop that watches the contract and acts on it Platform team
Child resources Namespaces, role bindings, quotas, workloads created inside the cluster Controller, with owner metadata
External infrastructure Cloud or appliance resources that Kubernetes cannot manage directly Controller, through the external provider’s API

A custom resource on its own only stores and returns structured data. Nothing is provisioned until a controller watches that data and acts on it, so the contract and the controller should be designed together.

Step 1: Define the service contract

Start with what a tenant should be able to declare. Typical candidates are a managed cluster, a namespace bundle with quotas and policies, an application environment, or an instance of a shared service. Choose one object for the first version. A service that tries to model everything at once usually ends up with a schema nobody wants to change.

Split the object into two halves. The spec holds desired state, written by tenants or the platform team. The status holds what the controller has observed. Kubernetes API conventions commonly express status as a list of conditions, each with a type, status, reason, message, and last transition time. Pick condition types that tenants can act on, such as a Ready condition for the whole object and a Progressing or Degraded condition when work is underway or stuck.

Describe the schema with an OpenAPI v3 structural schema in the CRD so that the API server rejects malformed input before your controller sees it. The object below is illustrative; the group, kind, and fields are placeholders for your own product decisions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
apiVersion: platform.example.com/v1alpha1
kind: TenantEnvironment
metadata:
  name: payments-staging
  namespace: tenant-payments
spec:
  tier: standard
  cpu: 8
  memory: 16Gi
  exposure: internal
status:
  observedGeneration: 3
  conditions:
  - type: Ready
    status: "False"
    reason: Provisioning
    message: child namespace not yet labeled

Step 2: Choose the API extension mechanism

Kubernetes offers two distinct ways to add an API. A CustomResourceDefinition (CRD) defines a new resource type that the Kubernetes control plane serves and stores. API aggregation registers a separately implemented extension API server, and the main API server proxies requests for the registered paths to it. The official extension documentation presents these as different mechanisms, so the choice is an operational decision rather than a naming preference.

Axis CRD API aggregation
What you provide A schema for a resource type that the control plane serves and stores A separately implemented API server registered with the aggregation layer
Storage Handled by the Kubernetes control plane Determined by the extension API server
Behavior beyond schema Standard API machinery plus schema validation Custom API semantics implemented in your server
Operational burden Controller and CRD; the control plane does the serving You run, scale, and secure an extra API server service
kubectl access Available once the CRD is installed Available for registered paths once the API service is registered
Authentication and authorization Uses the API server’s authentication, authorization, and audit logging Not stated in the Kubernetes extension documentation; confirm how your server delegates these checks
Best fit Schema-defined objects that controllers reconcile Specialized API behavior that a schema cannot express

For most tenant-facing platform objects, a CRD is the right starting point. Move to aggregation only when you can name a specific behavior that justifies running another API server with its own availability, scaling, and security configuration.

Step 3: Reconcile toward the declared state

A controller is a control loop. The Kubernetes controller documentation describes the idea this way: “In robotics and automation, a control loop is a non-terminating loop that regulates the state of a system.” For a platform service, the loop runs continuously, watches your custom resources, and works toward the state each one declares. Each pass should follow the same sequence:

  1. Read the object’s spec, status, and metadata, including metadata.generation, which identifies the version of the spec being processed.
  2. Read the child objects the controller owns and, where the service has them, the state of external infrastructure.
  3. Compare desired state with observed state and compute the changes needed.
  4. Apply those changes idempotently: create what is missing, update what has drifted, and delete what the spec no longer requires.
  5. Write status only when it has actually changed, with conditions that explain what the controller sees.
  6. Return success when the object has converged, requeue with backoff after transient errors, and record a failure reason the tenant can read.

Design every step to be safe to repeat. A pass can stop halfway, and the cluster can change between passes. The controller should not assume that the system ever reaches a stable final state, so each pass must work from what it observes now.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ownership of child resources

Controllers commonly create and update resources through the API server, such as namespaces, role bindings, quotas, and workloads. Set an owner reference from each child to its parent so that Kubernetes garbage collection removes children when the parent is deleted. Kubernetes also notes that several controllers can create the same kind of object, and ownership metadata is how each one distinguishes the objects it manages.

Give each controller one coherent responsibility. A controller that provisions tenant namespaces and a separate one that manages network exposure let you fail, upgrade, and debug each concern without blocking the other.

Deletion and external cleanup

Garbage collection handles in-cluster children. External infrastructure is different, because nothing in the cluster knows how to delete a cloud resource. When the controller first provisions external state, add a finalizer to the parent. The finalizer blocks deletion until the controller has torn down the external resources and removed itself from the list. That is what makes deletion safe, and it is also a common cause of stuck deletions, covered in the troubleshooting table below.

Step 4: Set tenancy and authorization as product requirements

The title does not determine how tenants share infrastructure, so this is a decision you must make and document. Namespace separation on its own is not a security boundary. Combine it with the RBAC, quota, and data-plane controls described below, and check the result against the threat model the service actually faces.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the tenancy model

Model Isolation Shared components Per-tenant overhead (generally) Controller implication
Shared cluster, namespace per tenant Namespaces, RBAC, quotas, and network policy, combined with other controls Control plane, nodes, and cluster-scoped add-ons Lowest One controller set serves all tenants and must scope every action to the tenant namespace
Virtual control plane per tenant Each tenant gets its own API server; workloads may still share nodes Host cluster and its nodes Moderate Controllers often need to sync objects between the tenant API and the host cluster
Dedicated cluster per tenant Separation at the cluster level Little beyond the provider and fleet tooling Highest Each cluster needs a controller deployment, or a fleet tool that manages them, which multiplies operational work

Document the chosen model before writing the controller, because it determines where the controller runs and what it is allowed to touch.

Grant RBAC for the new resource

CRDs use the API server’s authentication, authorization, and audit logging, so your objects share the same identities and audit trail as built-in resources. RBAC does not grant access to new resource types automatically, though. Tenants and operators get nothing until you grant it explicitly. The namespaced Role below lets tenant editors manage their objects. Bind it with a RoleBinding in each tenant namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: tenant-environment-editor
  namespace: tenant-payments
rules:
- apiGroups: [platform.example.com]
  resources: [tenantenvironments]
  verbs: [get, list, watch, create, update, patch, delete]

The Role names the main resource, not the tenantenvironments/status subresource. That keeps status writes with the controller. The example assumes the CRD is namespaced; a cluster-scoped CRD needs a ClusterRole and a different binding pattern.

Limit the controller and the tenant workloads

Give the controller’s ServiceAccount only the verbs it needs in the namespaces it manages, then verify the result before deploying. This command checks whether a hypothetical controller identity can create RoleBindings in a tenant namespace:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
kubectl auth can-i create rolebindings -n tenant-payments --as=system:serviceaccount:platform-system:env-controller

Set resource requests and limits for tenant workloads with ResourceQuota and LimitRange objects in each tenant namespace. Enforce data-plane isolation with measures suited to the service, such as NetworkPolicy and separate node pools where the threat model requires them. Kubernetes multi-tenancy guidance treats namespace handling, resource requests and limits, and data-plane isolation as explicit design items for operators; none of them follow automatically from creating a namespace.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Step 5: Make network ownership explicit

If the service exposes application traffic, decide who owns each layer before writing the controller. Gateway API models this with three roles. Its resources are implemented by controllers, which may provision a cloud load balancer or run an in-cluster proxy.

Role Responsibility in Gateway API How it maps to your service
Infrastructure provider Supplies the infrastructure that implements Gateway resources Usually invisible to tenants; provider-managed
Cluster operator Owns policy and network access, such as which gateways exist and which namespaces may attach to them Platform-team configuration, created by your controllers or by operators
Application developer Configures routes and service composition for an application Routes declared by tenants in their own namespaces

Before promising a behavior to tenants, confirm which Gateway API version and features your chosen implementation supports on your target cluster and provider. A resource the API server accepts may still not be programmed by any controller, so tenants need status they can read.

Step 6: Choose an implementation framework

The official Kubernetes documentation lists several community tools for writing operators. Among them are Kubebuilder, Operator SDK, Kopf, and Java Operator SDK, along with others. Kubebuilder and Operator SDK are commonly used with Go, Operator SDK also supports Ansible and Helm, Kopf is a Python framework, and Java Operator SDK targets Java. Inclusion in the documentation list is not an endorsement, and it does not indicate which version is current. Evaluate candidates on:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Language fit: can your team operate and debug the chosen language well?
  • Maintenance status: recent releases, responsive issue handling, and clear compatibility notes.
  • Generated API conventions: scaffolding for CRDs, generated helper code, and the webhooks you will need.
  • Testing support: whether you can exercise reconciliation logic against a real API server.
  • Kubernetes compatibility: which Kubernetes releases each framework and its dependencies target.

Failure modes and recovery

Most incidents in this design come from ownership, deletion, or authorization rather than from the loop itself. The table lists the common symptoms, the first check to run, and the recovery step.

Symptom Likely cause First check Recovery
Tenant object stuck in deletion A finalizer whose external cleanup is failing, or whose controller is down Read metadata.finalizers with kubectl get tenantenvironment payments-staging -n tenant-payments -o yaml Restore the controller first. Remove a finalizer by hand only after confirming the external resources are gone, or you will leave orphans behind
Tenant receives Forbidden on the new resource No RBAC rule or RoleBinding covers the resource kubectl auth can-i list tenantenvironments -n tenant-payments --as=<tenant-user> Add the Role and RoleBinding described in Step 4
Child objects keep changing Two controllers manage the same field, or a controller and a human both write it kubectl get deployment <name> -n <namespace> -o yaml --show-managed-fields Assign one manager per field and reconcile ownership metadata
API server receives a constant stream of updates The controller writes status on every pass, even when nothing changed Count status writes in controller logs Compare the new status with the stored one and write only real changes
Controller cannot create children in a tenant namespace The controller’s ServiceAccount lacks a Role for that verb and resource kubectl auth can-i with --as=system:serviceaccount:<namespace>:<name> Grant the narrowest role that covers the namespaces it manages
Gateway route is accepted but traffic does not flow No controller has programmed the resource, or the feature is unsupported Read the Gateway and route status conditions Confirm the implementation supports the features in use on your cluster

Upgrades and deletion risks

  • Deleting a CRD removes every custom resource of that type in the cluster. Restrict who can delete CRDs and place their changes under change control.
  • Schema changes that add a version need a planned conversion path. Settle the storage version and migration before changing the shape of objects tenants already use.
  • Test controller upgrades against objects created by the previous release. The new code must reconcile existing children without recreating them.

Version and provider limits

  • Feature availability depends on the Kubernetes version and on how your managed service is configured. Check both before relying on an extension point.
  • The Kubernetes extension documentation explains the mechanisms but does not establish what a particular provider supports. Confirm with your provider whether aggregated API servers and admission or conversion webhooks are allowed in your clusters.
  • Uptime, billing, and service-level commitments come from the provider’s own terms, not from Kubernetes documentation, so set your service objectives only after reading them.

The Bottom Line

Start with the smallest contract a tenant needs: one CRD with a status built from conditions, one controller responsible for one kind of child, and RBAC granted explicitly for the new resource. Decide the tenancy model before writing controller code, since it determines where the controller runs and what it may touch. Reach for API aggregation only when a specific requirement cannot be expressed as a CRD.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.