Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To run LiteLLM as a shared gateway in production, deploy the proxy as stateless services behind an HTTPS load balancer with at least two replicas. Use PostgreSQL for keys, teams, users, spend logs and configuration. Use Redis for rate limits, router state and caching shared across instances. Keep the master key and salt key in a secret manager, never in source control. Choose monolithic mode unless you need to scale the gateway, management APIs and UI independently.

This guide follows LiteLLM’s published production guidance, including its Production Deployment guide, and explains the reasoning behind each part of the setup.

Choose a deployment mode and platform path

LiteLLM documents two production modes. A monolithic deployment runs gateway traffic, management APIs and the UI in one service. A microservices deployment separates the gateway, the backend and the UI so each can be scaled on its own. The same guide documents platform paths: Helm charts for Amazon EKS, Google GKE and Azure AKS, and Terraform modules for AWS and Google Cloud.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option When it fits Trade-offs
Monolithic You want the simpler mode to operate, which LiteLLM describes as its simplest option. Gateway traffic, management APIs and the UI share one deployment, so they scale together.
Microservices You need to scale the gateway, backend and UI independently. More components to deploy, monitor and upgrade. Each role runs on its own service port.
Kubernetes with Helm You already operate EKS, GKE or AKS. You also manage the cluster, ingress, PostgreSQL, Redis and migrations.
Terraform modules You want documented infrastructure provisioning on AWS or Google Cloud. The guide lists modules for AWS and Google Cloud only. It has no Azure Terraform module and points Azure users to AKS with Helm.

This table compares documented deployment paths and their trade-offs. It is not a performance comparison, and LiteLLM’s documentation does not establish that either mode is faster or more reliable than the other.

Monolithic or microservices

Start with monolithic unless you have a measured reason to split. Every extra service adds a deployment target, a set of health checks and an upgrade path to manage. Microservices makes sense when the UI or management traffic has a different load profile from model traffic, or when you need the gateway pool to grow without also growing the UI.

Kubernetes, Helm or Terraform

Pick the path that matches the platform your team already runs. If you operate EKS, GKE or AKS, the Helm path fits your existing tooling. If you want infrastructure defined as code on AWS or Google Cloud, the Terraform modules are the documented starting point. Azure teams should plan on AKS with Helm, because the guide does not list an Azure Terraform module.

Production architecture and the roles of PostgreSQL and Redis

The documented layout places clients such as OpenAI SDK applications, LangChain applications or plain curl callers behind an HTTPS load balancer. Behind the load balancer sit two or more stateless LiteLLM replicas, with PostgreSQL and Redis as supporting services. Because the replicas do not keep state that must survive a restart, you add capacity by adding replicas.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PostgreSQL: keys, teams, users, spend and configuration

PostgreSQL stores keys, teams, users, spend logs and configuration. The proxy’s authentication and tracking features depend on it. If you need persistent virtual keys or spend history, you need a database. The consequences of running without one are covered in the spend-control section below.

Redis: shared state across instances

Redis supports rate limiting, router state and caching across instances. Without a shared Redis, rate limits, budgets and router cooldowns are counted separately inside each process. A limit meant for the whole cluster then becomes a per-process limit, so the effective ceiling rises with every replica you add. Any deployment with more than one replica should have Redis in place before it takes production traffic.

Migrations: one job, not every replica

The documented pattern runs schema changes in a migrations job once per upgrade. When that job owns migrations, the proxy instances should have schema updates disabled, so replicas do not each try to alter the database during a rollout. Check the production guide for the exact setting that controls this in your version.

Deployment sequence

  1. Provision PostgreSQL and Redis, and confirm that the proxy workloads can reach both.
  2. Generate the master key and salt key, then store them in your secret manager and inject them into the workloads. Do not place them in configuration files that live in source control.
  3. Run the migrations job once for the target version.
  4. Start two or more proxy replicas with schema updates disabled.
  5. Place the replicas behind an HTTPS load balancer, and deploy an image pinned to a version tag.
  6. Verify the path end to end. The quickstart does this with a local Docker Compose walkthrough: set up a model, create a virtual key and send one API request. Use a test key first.
  7. Enable metrics scraping and alerting, as described in the monitoring section.

Gateway credentials: master key and salt key

LiteLLM’s deployment has two secrets that behave differently, and both need deliberate handling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The master key

The master key authorizes management API operations. By default, it also serves as the Admin UI password. LiteLLM’s quickstart puts it this way:

“Anyone holding it has full admin access, so treat it like a root password, keep it out of source control, and rotate it if it ever leaks.”

The quote refers to LITELLM_MASTER_KEY. Source: LiteLLM documentation, Quickstart.

The salt key

The salt key encrypts provider API credentials that are persisted in the database. Generate it securely and record it in your secret manager before the first provider credential is saved. Do not change it after credentials have been stored. Doing so makes those stored provider credentials unreadable. Treat the salt key as a secret that must survive every database restore, and confirm that your backup process keeps the salt key alongside the database.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Spend control and budgets without a database

A process with no database can still expose an OpenAI-compatible API, which is useful for a quick trial. The quickstart limits what that mode can do, and the difference matters if you plan to enforce spend.

Capability Without a database With PostgreSQL
Admin UI model management Not available Available
Virtual keys Not available; virtual keys require a database Available
Spend tracking Not available; global spend remains unknown Available, with spend logs stored in PostgreSQL
Global budget enforcement A configured global budget does not stop requests Enforced only when spend is loaded from the database

If spend limits are a requirement, use the database-backed path. Provider-side spending limits can serve as an additional boundary outside the gateway, but they do not replace the gateway’s own tracking.

Attributing usage to keys with overwrite_user_with_key_hash

LiteLLM documents an optional overwrite_user_with_key_hash setting. When it is enabled, validated virtual-key and master-key requests have any caller-supplied user field replaced with a stable identity derived from the key. This gives you consistent attribution inside the gateway. Whether a provider transmits or maps that field is provider-dependent, so check the provider’s behavior before you rely on it for provider-side reporting.

Monitoring and alerting

Prometheus metrics and authentication

LiteLLM documents Prometheus metrics and Kubernetes autoscaling driven by request-rate or token-rate metrics. The main metrics endpoint sits behind virtual-key authentication. If your Prometheus server scrapes without credentials, you need a dedicated metrics listener. Use the official chart guidance for the exact metrics configuration in the deployment path you selected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Alerts to configure

LiteLLM’s production best-practices page describes alerts for:

  • model exceptions
  • slow or hanging requests
  • budget crossings
  • database errors
  • outages
  • spend reports

Observability integrations

The project overview names Langfuse, MLflow and Helicone among its observability callback integrations. Before you pick one, evaluate it against your requirements for traces, data retention, access controls and cost. Each integration sends request data to a third party, so the access and retention rules for that data belong in the decision.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Security and upgrade hygiene

A gateway handles provider credentials and request traffic, so the versions and artifacts you deploy are part of your security posture.

Package and image provenance after the March 2026 incident

According to the project’s own incident account in a GitHub issue, PyPI versions 1.82.7 and 1.82.8 were malicious in a March 2026 supply-chain incident. The same account reports that Docker image users were not affected. Treat this as the project’s statement about that event, not as a guarantee about every artifact or release. Install from official, signed container images at pinned version tags, and avoid unpinned installs from package indexes in production.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Advisory fixes: CVE-2026-42208 and CVE-2026-42271

Two official advisories name 1.83.7 as the patched release for their specific issues:

Advisory Affected versions Patched release named in advisory
CVE-2026-42208 1.81.16 and above, and below 1.83.7 1.83.7
CVE-2026-42271 Below 1.83.7 1.83.7

These statements cover only the two advisories above. They do not show that 1.83.7 is the latest recommended release, and they do not cover advisories published after these fixes.

Which version to deploy

This guide does not establish the current recommended LiteLLM release. Check the project’s release notes and its published security advisory list on the day you deploy, then pin the version you have tested in staging.

Container and proxy settings

Use signed official images with version tags, not a moving latest tag. Where relevant, configure trusted proxy ranges so the gateway accepts forwarding headers only from your load balancer. Apply schema changes through the documented migrations workflow rather than by hand.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Go-live checklist

  • Master key and salt key are stored in a secret manager, with rotation and backup procedures written down.
  • Images are pinned to version tags, and the tag has been checked against the current advisory list.
  • Migrations run as a separate job, and proxy replicas have schema updates disabled.
  • Redis is deployed for every multi-replica setup.
  • Metrics scraping works with the authentication model you chose.
  • Alerts exist for model exceptions, slow requests, budget crossings, database errors, outages and spend reports.
  • Spend limits are enforced through the database-backed path, with provider-side limits as an outer boundary where spend matters.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.