LiteLLM is an open-source AI gateway that proxies all LLM requests through a single endpoint, providing unified monitoring, access control, and cost tracking across providers like OpenAI and Anthropic. The post walks through running LiteLLM locally with Docker and PostgreSQL, configuring it via YAML, writing a demo Python client, integrating OpenTelemetry traces with VictoriaTraces, and exploring access management features including teams, users, API keys, budgets, and rate limits. Key span attributes for cost, token usage, model info, and user identity are highlighted. The author notes that while the monitoring and tracing capabilities are impressive, user management has some confusing behaviors — particularly around keys created outside a team bypassing team-level rate limits.
Table of contents
ContentsLiteLLM — main featuresWhy do we need this?Running LiteLLM with DockerConfig.yaml — LiteLLM configurationDemo Python App — AI ClientMonitoring, OpenTelemetry and TracesOpenTelemetry and VictoriaTracesGet Arseny Zinchenko (setevoy)’s stories in your inboxLiteLLM Span AttributesAccess ManagementAuthentication and accessTeams and UsersBudgets and limitsRBAC and System RolesCreating a TeamCreating a User in the Web UITeam PermissionsCreating a User API Key for a Team in the Web UICreating a User API Key in the Web UI without a Team and without LimitsCreating a User and API Key via the LiteLLM API with a Rate LimitInstead of conclusionsQuestions this post answers
How do I send LiteLLM proxy traces to an OpenTelemetry backend like VictoriaTraces?
Enable the otel callback in litellm_settings.callbacks in the config, then set OTEL_EXPORTER_OTLP_TRACES_ENDPOINT and OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf (or the LiteLLM-specific OTEL_EXPORTER and OTEL_ENDPOINT variables) pointing at the backend's OTLP ingestion path, such as /insert/opentelemetry/v1/traces for VictoriaTraces. Multiple exporters like langfuse or arize can run simultaneously alongside otel. Teams wiring LLM observability into existing tracing stacks can track gateway configuration patterns like this on daily.dev.
Why does a LiteLLM Team RPM limit not stop a user from making unlimited requests?
Team rate limits in LiteLLM only apply to API keys explicitly created for that team (keys carrying a team_id). A user with the internal-user role can still generate personal keys outside any team through the Web UI, and those keys have no rate limit unless upperbound_key_generate_params is configured, effectively letting them bypass team-level RPM and budget restrictions entirely. Anyone hardening multi-tenant LLM access controls can compare these gaps before rolling out gateway permissions on daily.dev.
How do I set a per-user rate limit when creating a LiteLLM user through the API instead of the Web UI?
Call POST /user/new with a JSON body including rpm_limit, user_role, and user_email; the response returns a generated API key already scoped to that limit. This works even though the Web UI at the time only exposes model access restrictions when creating a user, not budget or RPM/TPM fields, making the API the more complete option for enforcing per-user limits. Developers scripting LLM gateway provisioning can weigh API-versus-UI gaps like this via daily.dev.
Share this post