Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To run LiteLLM as a shared gateway in production, deploy the proxy as stateless services behind an HTTPS load balancer with at least two replicas. Use PostgreSQL for keys, teams, users, spend logs and configuration. Use Redis for rate limits, router state and caching shared across instances. Keep the master key and salt key in a secret manager, never in source control. Choose monolithic mode unless you need to scale the gateway, management APIs and UI independently.
This guide follows LiteLLM’s published production guidance, including its Production Deployment guide, and explains the reasoning behind each part of the setup.
Choose a deployment mode and platform path
LiteLLM documents two production modes. A monolithic deployment runs gateway traffic, management APIs and the UI in one service. A microservices deployment separates the gateway, the backend and the UI so each can be scaled on its own. The same guide documents platform paths: Helm charts for Amazon EKS, Google GKE and Azure AKS, and Terraform modules for AWS and Google Cloud.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Option | When it fits | Trade-offs |
|---|---|---|
| Monolithic | You want the simpler mode to operate, which LiteLLM describes as its simplest option. | Gateway traffic, management APIs and the UI share one deployment, so they scale together. |
| Microservices | You need to scale the gateway, backend and UI independently. | More components to deploy, monitor and upgrade. Each role runs on its own service port. |
| Kubernetes with Helm | You already operate EKS, GKE or AKS. | You also manage the cluster, ingress, PostgreSQL, Redis and migrations. |
| Terraform modules | You want documented infrastructure provisioning on AWS or Google Cloud. | The guide lists modules for AWS and Google Cloud only. It has no Azure Terraform module and points Azure users to AKS with Helm. |
This table compares documented deployment paths and their trade-offs. It is not a performance comparison, and LiteLLM’s documentation does not establish that either mode is faster or more reliable than the other.
#1 Best Overall
Monolithic or microservices
Start with monolithic unless you have a measured reason to split. Every extra service adds a deployment target, a set of health checks and an upgrade path to manage. Microservices makes sense when the UI or management traffic has a different load profile from model traffic, or when you need the gateway pool to grow without also growing the UI.
Kubernetes, Helm or Terraform
Pick the path that matches the platform your team already runs. If you operate EKS, GKE or AKS, the Helm path fits your existing tooling. If you want infrastructure defined as code on AWS or Google Cloud, the Terraform modules are the documented starting point. Azure teams should plan on AKS with Helm, because the guide does not list an Azure Terraform module.
Production architecture and the roles of PostgreSQL and Redis
The documented layout places clients such as OpenAI SDK applications, LangChain applications or plain curl callers behind an HTTPS load balancer. Behind the load balancer sit two or more stateless LiteLLM replicas, with PostgreSQL and Redis as supporting services. Because the replicas do not keep state that must survive a restart, you add capacity by adding replicas.
PostgreSQL: keys, teams, users, spend and configuration
PostgreSQL stores keys, teams, users, spend logs and configuration. The proxy’s authentication and tracking features depend on it. If you need persistent virtual keys or spend history, you need a database. The consequences of running without one are covered in the spend-control section below.
Rank #2
Redis: shared state across instances
Redis supports rate limiting, router state and caching across instances. Without a shared Redis, rate limits, budgets and router cooldowns are counted separately inside each process. A limit meant for the whole cluster then becomes a per-process limit, so the effective ceiling rises with every replica you add. Any deployment with more than one replica should have Redis in place before it takes production traffic.
Migrations: one job, not every replica
The documented pattern runs schema changes in a migrations job once per upgrade. When that job owns migrations, the proxy instances should have schema updates disabled, so replicas do not each try to alter the database during a rollout. Check the production guide for the exact setting that controls this in your version.
Deployment sequence
- Provision PostgreSQL and Redis, and confirm that the proxy workloads can reach both.
- Generate the master key and salt key, then store them in your secret manager and inject them into the workloads. Do not place them in configuration files that live in source control.
- Run the migrations job once for the target version.
- Start two or more proxy replicas with schema updates disabled.
- Place the replicas behind an HTTPS load balancer, and deploy an image pinned to a version tag.
- Verify the path end to end. The quickstart does this with a local Docker Compose walkthrough: set up a model, create a virtual key and send one API request. Use a test key first.
- Enable metrics scraping and alerting, as described in the monitoring section.
Gateway credentials: master key and salt key
LiteLLM’s deployment has two secrets that behave differently, and both need deliberate handling.
The master key
The master key authorizes management API operations. By default, it also serves as the Admin UI password. LiteLLM’s quickstart puts it this way:
Rank #3
“Anyone holding it has full admin access, so treat it like a root password, keep it out of source control, and rotate it if it ever leaks.”
The quote refers to LITELLM_MASTER_KEY. Source: LiteLLM documentation, Quickstart.
The salt key
The salt key encrypts provider API credentials that are persisted in the database. Generate it securely and record it in your secret manager before the first provider credential is saved. Do not change it after credentials have been stored. Doing so makes those stored provider credentials unreadable. Treat the salt key as a secret that must survive every database restore, and confirm that your backup process keeps the salt key alongside the database.
Spend control and budgets without a database
A process with no database can still expose an OpenAI-compatible API, which is useful for a quick trial. The quickstart limits what that mode can do, and the difference matters if you plan to enforce spend.
| Capability | Without a database | With PostgreSQL |
|---|---|---|
| Admin UI model management | Not available | Available |
| Virtual keys | Not available; virtual keys require a database | Available |
| Spend tracking | Not available; global spend remains unknown | Available, with spend logs stored in PostgreSQL |
| Global budget enforcement | A configured global budget does not stop requests | Enforced only when spend is loaded from the database |
If spend limits are a requirement, use the database-backed path. Provider-side spending limits can serve as an additional boundary outside the gateway, but they do not replace the gateway’s own tracking.
Attributing usage to keys with overwrite_user_with_key_hash
LiteLLM documents an optional overwrite_user_with_key_hash setting. When it is enabled, validated virtual-key and master-key requests have any caller-supplied user field replaced with a stable identity derived from the key. This gives you consistent attribution inside the gateway. Whether a provider transmits or maps that field is provider-dependent, so check the provider’s behavior before you rely on it for provider-side reporting.
Monitoring and alerting
Prometheus metrics and authentication
LiteLLM documents Prometheus metrics and Kubernetes autoscaling driven by request-rate or token-rate metrics. The main metrics endpoint sits behind virtual-key authentication. If your Prometheus server scrapes without credentials, you need a dedicated metrics listener. Use the official chart guidance for the exact metrics configuration in the deployment path you selected.
Alerts to configure
LiteLLM’s production best-practices page describes alerts for:
- model exceptions
- slow or hanging requests
- budget crossings
- database errors
- outages
- spend reports
Observability integrations
The project overview names Langfuse, MLflow and Helicone among its observability callback integrations. Before you pick one, evaluate it against your requirements for traces, data retention, access controls and cost. Each integration sends request data to a third party, so the access and retention rules for that data belong in the decision.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Security and upgrade hygiene
A gateway handles provider credentials and request traffic, so the versions and artifacts you deploy are part of your security posture.
Package and image provenance after the March 2026 incident
According to the project’s own incident account in a GitHub issue, PyPI versions 1.82.7 and 1.82.8 were malicious in a March 2026 supply-chain incident. The same account reports that Docker image users were not affected. Treat this as the project’s statement about that event, not as a guarantee about every artifact or release. Install from official, signed container images at pinned version tags, and avoid unpinned installs from package indexes in production.
Free tools Windows power users keep installed
One-click scans. No signup required.
Advisory fixes: CVE-2026-42208 and CVE-2026-42271
Two official advisories name 1.83.7 as the patched release for their specific issues:
| Advisory | Affected versions | Patched release named in advisory |
|---|---|---|
| CVE-2026-42208 | 1.81.16 and above, and below 1.83.7 | 1.83.7 |
| CVE-2026-42271 | Below 1.83.7 | 1.83.7 |
These statements cover only the two advisories above. They do not show that 1.83.7 is the latest recommended release, and they do not cover advisories published after these fixes.
Which version to deploy
This guide does not establish the current recommended LiteLLM release. Check the project’s release notes and its published security advisory list on the day you deploy, then pin the version you have tested in staging.
Container and proxy settings
Use signed official images with version tags, not a moving latest tag. Where relevant, configure trusted proxy ranges so the gateway accepts forwarding headers only from your load balancer. Apply schema changes through the documented migrations workflow rather than by hand.
Recommended Free Tools
Quick Recap
Go-live checklist
- Master key and salt key are stored in a secret manager, with rotation and backup procedures written down.
- Images are pinned to version tags, and the tag has been checked against the current advisory list.
- Migrations run as a separate job, and proxy replicas have schema updates disabled.
- Redis is deployed for every multi-replica setup.
- Metrics scraping works with the authentication model you chose.
- Alerts exist for model exceptions, slow requests, budget crossings, database errors, outages and spend reports.
- Spend limits are enforced through the database-backed path, with provider-side limits as an outer boundary where spend matters.
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

