Deploy the gateway between your applications and AI providers, then give each application, user, or team a narrowly scoped virtual key. For a first deployment, LiteLLM’s Docker quickstart provides a gateway on port 4000 and a PostgreSQL-backed management workflow. For production, use a database-backed, monitored deployment with network isolation and protected secrets; a configured budget is not a reliable spending cap in LiteLLM’s DB-less mode.
Choose a deployment shape
The right shape depends on whether you are learning the request path or operating a shared service. LiteLLM is a concrete example, not the only way to build a self-hosted inference gateway. Its documentation, accessed October 4, 2026, describes a Docker quickstart and production options ranging from a monolithic service to separated microservices. Check the documentation for the exact release you deploy: labels, configuration, and behavior can change.
| Deployment shape | Good fit | Database and scaling | Operational ownership |
|---|---|---|---|
| Single-machine Docker quickstart | Learning the request path or evaluating a gateway on one host. | The documented sample starts a gateway and PostgreSQL-backed workflow. It is not a multi-replica architecture. | You operate Docker Compose, protect generated secrets, configure a model, and issue virtual keys. See LiteLLM’s Docker quickstart. |
| Production monolithic service | A shared deployment where one service shape is simpler to operate. | Use PostgreSQL for keys, teams, users, configuration, and spend logs; use Redis for shared rate limiting, router state, and caching when running multiple instances. | You own ingress, secret storage, TLS, migrations, monitoring, and scaling. The production deployment guide describes this approach. |
| Production microservices | Deployments that need inference traffic to scale independently from management APIs and the UI. | The guide separates gateway traffic, management backend, and UI; production support services still include PostgreSQL and Redis for the documented multi-instance setup. | More components mean more deployment and monitoring responsibilities. The production guide describes Helm on EKS, GKE, or AKS and official Terraform modules for AWS and GCP. |
A single-machine setup can help validate configuration, but it does not by itself provide production isolation, availability, or multi-instance coordination. Choose monolithic versus microservices based on the need to scale inference separately, not on an assumption that the more distributed design is automatically safer.
Bring up the gateway and test the request path
In LiteLLM’s documented Docker quickstart, the gateway listens on port 4000. The sample flow generates a master key and salt key before starting Docker Compose, connects a model, creates a virtual key, and sends an OpenAI-compatible client request through the gateway. Follow the current quickstart for its exact commands and configuration rather than copying an unpinned example into production.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
- Review the Compose configuration. Use the official Docker quickstart and inspect the Compose file before starting services. Keep the master key, salt key, and provider credentials out of source control.
- Start the gateway and database-backed workflow. Confirm the service is reachable on the intended interface and port. The quickstart’s port 4000 is a sample deployment detail, not a reason to expose that port publicly.
- Connect a model. Configure the provider credentials on the server and define the models the proxy should expose to clients.
- Issue a virtual key with limited access. Set its allowed models and associate it with the right user or team before giving it to an application.
- Send a client request through the gateway. Use the virtual key as the client credential and verify that the request reaches only an allowed model. This tests the path applications will use instead of giving them raw provider credentials.
For repeatable production deployments, pin a released container image tag instead of relying blindly on a moving latest tag. Treat the sample secrets as secrets: generate and store production values securely, and do not commit them with application code.
Build a production topology around shared state
When multiple gateway replicas serve requests, the support services need clear jobs. In LiteLLM’s production guidance, PostgreSQL stores keys, teams, users, configuration, and spend logs. Redis supports rate limiting, router state, and caching across instances. A load balancer distributes traffic to stateless replicas.
Rank #2
- PostgreSQL: Use the connected database for persistent identity and spend-related records that management and budget enforcement depend on.
- Redis: Use shared Redis state for the documented multi-instance rate-limiting, router-state, and caching functions; do not assume separate replicas coordinate these features without shared state.
- Schema migration: Run a migration job during schema changes and upgrades. The production guide’s migration-job pattern says proxy instances should not independently run schema updates when that job is used.
- Ingress and secrets: Put replicas behind a load balancer and place provider and administrator secrets in server-side environment-backed or managed secret storage.
The official guide describes Kubernetes deployment through Helm for EKS, GKE, or AKS, as well as Terraform modules for AWS and GCP. These are documented deployment paths, not a requirement to use a particular cloud provider.
Assign roles, keys, and model access separately
Give each application, user, or team its own virtual key rather than distributing a provider key or a gateway administrator credential. Configure the key’s permitted models and attach the appropriate user or team relationship and limits. LiteLLM’s documentation says spend can be recorded at key, user, and team level when those identifiers are attached. See LiteLLM’s virtual-key documentation.
Recommended Free Tools
Rank #3
Model permissions and management permissions are distinct. LiteLLM evaluates model access at the key, but management routes for actions such as administering keys, users, and teams depend on the owning user’s role. A key associated with a proxy administrator may therefore have access to management endpoints even if it was created for an application. Where relevant, explicitly constrain routes with allowed_routes; do not infer that an application key is low privilege from its intended use alone.
For an organization-wide identity design, gateway virtual keys are one option, not a complete identity architecture. Amazon Web Services’ guidance describes an AWS-specific alternative using identity-based authentication with short-lived credentials and signed requests, including mapping end-user OAuth2/OIDC identities to roles in AWS environments. Select that pattern only when it fits the surrounding AWS identity and request-signing architecture; it is not a generic LiteLLM setting. See AWS, Generative AI inference architecture and best practices on AWS.
Rank #4
Set token-throughput limits and spend budgets for the right scope
Token quotas can mean different controls. TPM and RPM limits constrain token throughput or request rate. A max_budget is a spend limit over a configured period; adding a budget does not automatically impose TPM or RPM limits. Decide whether a policy belongs to an individual key, a user, a team, or the whole proxy, and define its reset period. If a team key should be constrained by both its own policy and the team’s policy, verify that inheritance behavior on the selected release.
For LiteLLM, budget enforcement depends on a connected database: the proxy compares requests against spend read from the database. The documentation states, “Budgets require a database.” In DB-less mode, global budget checks fail open, so the proxy can continue serving beyond the configured amount; key and team virtual-key budgets are unavailable too. A budget value in a configuration file is not evidence that a hard cap is being enforced.
Best Value
Before relying on a monetary limit, test the enforcement path in the deployed release with the database connected, the intended key and team associations in place, and representative request patterns. The documentation notes that some routes without token pricing enforce against recorded spend rather than a reserved cost estimate. That distinction matters if you expect the gateway to stop a request based on its predicted final cost rather than previously recorded usage.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Protect the endpoint, credentials, and operations
Authentication alone is not enough. Put the gateway behind HTTPS/TLS and restrict ingress to the systems that need it. Prefer a private or internal load balancer for company services; if public exposure is unavoidable, use IP restrictions where possible and appropriate edge protections. AWS’s guidance puts the reason plainly: “Network isolation complements API keys and identity-based authentication so that even leaked credentials can’t reach an endpoint directly.”
Quick Recap
- Keep secrets server-side. Do not embed provider API keys or gateway administrator secrets in browser code or application bundles. Store them in protected environment-backed configuration or managed secret storage.
- Rotate and scope credentials. Rotate short-term credentials and limit each key to the models and actions it needs. Do not use a master key as an ordinary application credential.
- Monitor the service. Track latency, throughput, errors, and resource use. The LiteLLM Kubernetes guidance describes metrics endpoints and autoscaling options; account for the fact that tokens-per-second signals for streaming requests are counted when a response completes.
- Exercise failures deliberately. Confirm that a disallowed model is rejected, budget checks work with the database connected, and a database or Redis outage has the expected effect on the features that depend on it. Do not claim a quota or shared limit is protective until its enforcement behavior has been observed on the chosen release.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

