iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Kubernetes as a service works when a tenant can declare what they need in one small API object, and a controller keeps the cluster and any external systems moving toward that declaration. The custom resource is the service contract. The controller is the reconciliation engine that creates and maintains the Kubernetes objects and external infrastructure behind it.
Most of the hard decisions follow from that split: which extension mechanism exposes the contract, how the controller reports progress, how tenants are isolated, who owns networking, and which framework you use to write the loop. This guide stays architectural. It does not assume a cloud provider, tenancy model, or service-level objective, and it flags where a choice depends on those.
How the pieces fit together
A platform of this kind has four parts. Keeping them separate in your design makes ownership and failure handling easier to reason about.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Layer | What it does | Typical owner |
|---|---|---|
| Service contract | A custom resource with a spec for desired state and a status for observed state |
Platform team designs it; tenants create instances |
| Controller | A control loop that watches the contract and acts on it | Platform team |
| Child resources | Namespaces, role bindings, quotas, workloads created inside the cluster | Controller, with owner metadata |
| External infrastructure | Cloud or appliance resources that Kubernetes cannot manage directly | Controller, through the external provider’s API |
A custom resource on its own only stores and returns structured data. Nothing is provisioned until a controller watches that data and acts on it, so the contract and the controller should be designed together.
#1 Best Overall
Step 1: Define the service contract
Start with what a tenant should be able to declare. Typical candidates are a managed cluster, a namespace bundle with quotas and policies, an application environment, or an instance of a shared service. Choose one object for the first version. A service that tries to model everything at once usually ends up with a schema nobody wants to change.
Split the object into two halves. The spec holds desired state, written by tenants or the platform team. The status holds what the controller has observed. Kubernetes API conventions commonly express status as a list of conditions, each with a type, status, reason, message, and last transition time. Pick condition types that tenants can act on, such as a Ready condition for the whole object and a Progressing or Degraded condition when work is underway or stuck.
Describe the schema with an OpenAPI v3 structural schema in the CRD so that the API server rejects malformed input before your controller sees it. The object below is illustrative; the group, kind, and fields are placeholders for your own product decisions:
Recommended Free Tools
apiVersion: platform.example.com/v1alpha1
kind: TenantEnvironment
metadata:
name: payments-staging
namespace: tenant-payments
spec:
tier: standard
cpu: 8
memory: 16Gi
exposure: internal
status:
observedGeneration: 3
conditions:
- type: Ready
status: "False"
reason: Provisioning
message: child namespace not yet labeled
Step 2: Choose the API extension mechanism
Kubernetes offers two distinct ways to add an API. A CustomResourceDefinition (CRD) defines a new resource type that the Kubernetes control plane serves and stores. API aggregation registers a separately implemented extension API server, and the main API server proxies requests for the registered paths to it. The official extension documentation presents these as different mechanisms, so the choice is an operational decision rather than a naming preference.
| Axis | CRD | API aggregation |
|---|---|---|
| What you provide | A schema for a resource type that the control plane serves and stores | A separately implemented API server registered with the aggregation layer |
| Storage | Handled by the Kubernetes control plane | Determined by the extension API server |
| Behavior beyond schema | Standard API machinery plus schema validation | Custom API semantics implemented in your server |
| Operational burden | Controller and CRD; the control plane does the serving | You run, scale, and secure an extra API server service |
| kubectl access | Available once the CRD is installed | Available for registered paths once the API service is registered |
| Authentication and authorization | Uses the API server’s authentication, authorization, and audit logging | Not stated in the Kubernetes extension documentation; confirm how your server delegates these checks |
| Best fit | Schema-defined objects that controllers reconcile | Specialized API behavior that a schema cannot express |
For most tenant-facing platform objects, a CRD is the right starting point. Move to aggregation only when you can name a specific behavior that justifies running another API server with its own availability, scaling, and security configuration.
Step 3: Reconcile toward the declared state
A controller is a control loop. The Kubernetes controller documentation describes the idea this way: “In robotics and automation, a control loop is a non-terminating loop that regulates the state of a system.” For a platform service, the loop runs continuously, watches your custom resources, and works toward the state each one declares. Each pass should follow the same sequence:
- Read the object’s
spec,status, andmetadata, includingmetadata.generation, which identifies the version of the spec being processed. - Read the child objects the controller owns and, where the service has them, the state of external infrastructure.
- Compare desired state with observed state and compute the changes needed.
- Apply those changes idempotently: create what is missing, update what has drifted, and delete what the spec no longer requires.
- Write status only when it has actually changed, with conditions that explain what the controller sees.
- Return success when the object has converged, requeue with backoff after transient errors, and record a failure reason the tenant can read.
Design every step to be safe to repeat. A pass can stop halfway, and the cluster can change between passes. The controller should not assume that the system ever reaches a stable final state, so each pass must work from what it observes now.
Ownership of child resources
Controllers commonly create and update resources through the API server, such as namespaces, role bindings, quotas, and workloads. Set an owner reference from each child to its parent so that Kubernetes garbage collection removes children when the parent is deleted. Kubernetes also notes that several controllers can create the same kind of object, and ownership metadata is how each one distinguishes the objects it manages.
Rank #3
Give each controller one coherent responsibility. A controller that provisions tenant namespaces and a separate one that manages network exposure let you fail, upgrade, and debug each concern without blocking the other.
Deletion and external cleanup
Garbage collection handles in-cluster children. External infrastructure is different, because nothing in the cluster knows how to delete a cloud resource. When the controller first provisions external state, add a finalizer to the parent. The finalizer blocks deletion until the controller has torn down the external resources and removed itself from the list. That is what makes deletion safe, and it is also a common cause of stuck deletions, covered in the troubleshooting table below.
Step 4: Set tenancy and authorization as product requirements
The title does not determine how tenants share infrastructure, so this is a decision you must make and document. Namespace separation on its own is not a security boundary. Combine it with the RBAC, quota, and data-plane controls described below, and check the result against the threat model the service actually faces.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Choose the tenancy model
| Model | Isolation | Shared components | Per-tenant overhead (generally) | Controller implication |
|---|---|---|---|---|
| Shared cluster, namespace per tenant | Namespaces, RBAC, quotas, and network policy, combined with other controls | Control plane, nodes, and cluster-scoped add-ons | Lowest | One controller set serves all tenants and must scope every action to the tenant namespace |
| Virtual control plane per tenant | Each tenant gets its own API server; workloads may still share nodes | Host cluster and its nodes | Moderate | Controllers often need to sync objects between the tenant API and the host cluster |
| Dedicated cluster per tenant | Separation at the cluster level | Little beyond the provider and fleet tooling | Highest | Each cluster needs a controller deployment, or a fleet tool that manages them, which multiplies operational work |
Document the chosen model before writing the controller, because it determines where the controller runs and what it is allowed to touch.
Grant RBAC for the new resource
CRDs use the API server’s authentication, authorization, and audit logging, so your objects share the same identities and audit trail as built-in resources. RBAC does not grant access to new resource types automatically, though. Tenants and operators get nothing until you grant it explicitly. The namespaced Role below lets tenant editors manage their objects. Bind it with a RoleBinding in each tenant namespace:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: tenant-environment-editor
namespace: tenant-payments
rules:
- apiGroups: [platform.example.com]
resources: [tenantenvironments]
verbs: [get, list, watch, create, update, patch, delete]
The Role names the main resource, not the tenantenvironments/status subresource. That keeps status writes with the controller. The example assumes the CRD is namespaced; a cluster-scoped CRD needs a ClusterRole and a different binding pattern.
Limit the controller and the tenant workloads
Give the controller’s ServiceAccount only the verbs it needs in the namespaces it manages, then verify the result before deploying. This command checks whether a hypothetical controller identity can create RoleBindings in a tenant namespace:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11kubectl auth can-i create rolebindings -n tenant-payments --as=system:serviceaccount:platform-system:env-controller
Set resource requests and limits for tenant workloads with ResourceQuota and LimitRange objects in each tenant namespace. Enforce data-plane isolation with measures suited to the service, such as NetworkPolicy and separate node pools where the threat model requires them. Kubernetes multi-tenancy guidance treats namespace handling, resource requests and limits, and data-plane isolation as explicit design items for operators; none of them follow automatically from creating a namespace.
Best Value
Step 5: Make network ownership explicit
If the service exposes application traffic, decide who owns each layer before writing the controller. Gateway API models this with three roles. Its resources are implemented by controllers, which may provision a cloud load balancer or run an in-cluster proxy.
| Role | Responsibility in Gateway API | How it maps to your service |
|---|---|---|
| Infrastructure provider | Supplies the infrastructure that implements Gateway resources | Usually invisible to tenants; provider-managed |
| Cluster operator | Owns policy and network access, such as which gateways exist and which namespaces may attach to them | Platform-team configuration, created by your controllers or by operators |
| Application developer | Configures routes and service composition for an application | Routes declared by tenants in their own namespaces |
Before promising a behavior to tenants, confirm which Gateway API version and features your chosen implementation supports on your target cluster and provider. A resource the API server accepts may still not be programmed by any controller, so tenants need status they can read.
Step 6: Choose an implementation framework
The official Kubernetes documentation lists several community tools for writing operators. Among them are Kubebuilder, Operator SDK, Kopf, and Java Operator SDK, along with others. Kubebuilder and Operator SDK are commonly used with Go, Operator SDK also supports Ansible and Helm, Kopf is a Python framework, and Java Operator SDK targets Java. Inclusion in the documentation list is not an endorsement, and it does not indicate which version is current. Evaluate candidates on:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →- Language fit: can your team operate and debug the chosen language well?
- Maintenance status: recent releases, responsive issue handling, and clear compatibility notes.
- Generated API conventions: scaffolding for CRDs, generated helper code, and the webhooks you will need.
- Testing support: whether you can exercise reconciliation logic against a real API server.
- Kubernetes compatibility: which Kubernetes releases each framework and its dependencies target.
Failure modes and recovery
Most incidents in this design come from ownership, deletion, or authorization rather than from the loop itself. The table lists the common symptoms, the first check to run, and the recovery step.
| Symptom | Likely cause | First check | Recovery |
|---|---|---|---|
| Tenant object stuck in deletion | A finalizer whose external cleanup is failing, or whose controller is down | Read metadata.finalizers with kubectl get tenantenvironment payments-staging -n tenant-payments -o yaml |
Restore the controller first. Remove a finalizer by hand only after confirming the external resources are gone, or you will leave orphans behind |
| Tenant receives Forbidden on the new resource | No RBAC rule or RoleBinding covers the resource | kubectl auth can-i list tenantenvironments -n tenant-payments --as=<tenant-user> |
Add the Role and RoleBinding described in Step 4 |
| Child objects keep changing | Two controllers manage the same field, or a controller and a human both write it | kubectl get deployment <name> -n <namespace> -o yaml --show-managed-fields |
Assign one manager per field and reconcile ownership metadata |
| API server receives a constant stream of updates | The controller writes status on every pass, even when nothing changed | Count status writes in controller logs | Compare the new status with the stored one and write only real changes |
| Controller cannot create children in a tenant namespace | The controller’s ServiceAccount lacks a Role for that verb and resource | kubectl auth can-i with --as=system:serviceaccount:<namespace>:<name> |
Grant the narrowest role that covers the namespaces it manages |
| Gateway route is accepted but traffic does not flow | No controller has programmed the resource, or the feature is unsupported | Read the Gateway and route status conditions | Confirm the implementation supports the features in use on your cluster |
Upgrades and deletion risks
- Deleting a CRD removes every custom resource of that type in the cluster. Restrict who can delete CRDs and place their changes under change control.
- Schema changes that add a version need a planned conversion path. Settle the storage version and migration before changing the shape of objects tenants already use.
- Test controller upgrades against objects created by the previous release. The new code must reconcile existing children without recreating them.
Version and provider limits
- Feature availability depends on the Kubernetes version and on how your managed service is configured. Check both before relying on an extension point.
- The Kubernetes extension documentation explains the mechanisms but does not establish what a particular provider supports. Confirm with your provider whether aggregated API servers and admission or conversion webhooks are allowed in your clusters.
- Uptime, billing, and service-level commitments come from the provider’s own terms, not from Kubernetes documentation, so set your service objectives only after reading them.
The Bottom Line
Start with the smallest contract a tenant needs: one CRD with a status built from conditions, one controller responsible for one kind of child, and RBAC granted explicitly for the new resource. Decide the tenancy model before writing controller code, since it determines where the controller runs and what it may touch. Reach for API aggregation only when a specific requirement cannot be expressed as a CRD.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

