- Remove duplicate description kwarg in supported_db_objects Field()
that caused SyntaxError preventing proxy startup
- Wrap sync _get_request_headers() in asyncio.to_thread to avoid
blocking the event loop during AppRole/TLS cert auth
- Include exception messages in error responses for admin-only
endpoints to aid debugging
- approle_role_id is a non-secret identifier (like a username) per
Vault's AppRole model; masking it hinders admin auditing
- Use async httpx client for the token lookup-self call to avoid
blocking the event loop
Without this, deployments using supported_db_objects filtering would
silently skip polling for config_overrides, preventing Hashicorp Vault
config from syncing across pods.
When the messages or response JSON fields in spend logs are truncated
before being written to the database, the truncation marker now includes
a note explaining:
- This is a DB storage safeguard
- Full, untruncated data is still sent to logging callbacks (OTEL, Datadog, etc.)
- The MAX_STRING_LENGTH_PROMPT_IN_DB env var can be used to increase the limit
Also emits a verbose_proxy_logger.info message when truncation occurs in
the request body or response spend log paths.
Adds 3 new tests:
- test_truncation_includes_db_safeguard_note
- test_response_truncation_logs_info_message
- test_request_body_truncation_logs_info_message
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Add CRUD endpoints for managing Hashicorp Vault configuration via the
proxy admin API, with background sync, env var management, and
connection testing. Fix pre-existing bug where premium check ran after
global state mutation, and guard DELETE against clearing non-Vault
secret managers.
After a server restart, RBAC flags (disable_agents_for_internal_users, etc.)
were only loaded into general_settings lazily when get_ui_settings or
update_ui_settings was called. This meant restrictions were silently bypassed
until someone hit one of those endpoints.
Add _sync_ui_settings_to_general_settings() to ProxyStartupEvent that reads
the litellm_uisettings DB record and syncs _RUNTIME_GENERAL_SETTINGS_FLAGS
into general_settings immediately after prisma_client is initialized.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Previously, model_dump(exclude_none=True) included all bool fields (since
False != None), causing a partial PATCH to overwrite every other setting to
its default. Fix uses exclude_unset=True and reads the existing DB record
before merging, giving proper PATCH semantics.
This was a pre-existing bug but is fixed here since we're touching this code.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- rbac_utils.py: change feature_name from str to Literal["agents", "vector_stores"]
so typos are caught by type checkers at import time
- proxy_setting_endpoints.py: extract _RUNTIME_GENERAL_SETTINGS_FLAGS as a module-level
constant, replacing duplicated inline lists in get_ui_settings and update_ui_settings
- test_vector_store_rbac.py: remove try/except pattern that silently swallowed non-403
HTTPExceptions; tests now let any unexpected exception propagate as a test failure
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- rbac_utils.py: remove duplicated _check_if_team_admin/_is_user_team_admin_for_any_team;
delegate to _user_has_admin_privileges from management_endpoints/common_utils with the
shared user_api_key_cache (fixes no-op DualCache and missing org admin coverage)
- test_rbac_utils.py: update patch target to match new delegation path
- SidebarProvider.tsx: pass allowAgentsForTeamAdmins and allowVectorStoresForTeamAdmins
props to Sidebar
- leftnav.tsx: add useTeams hook + isTeamAdmin memo; exempt team admins from sidebar
filtering when allow_*_for_team_admins is enabled (fixes frontend/backend inconsistency)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add 58 unit tests across 5 files in ui/litellm-dashboard/src/components/policies/:
- build_attachment_data.test.ts: pure logic tests for global vs specific scope
- impact_preview_alert.test.tsx: rendering for global scope warning and specific scope counts/tags
- PolicySelector.test.tsx: pure function tests for getPolicyOptionEntries/policyVersionRef + component fetch/disable behavior
- policy_info.test.tsx: loading, not-found, policy details rendering, admin gating, onClose/onEdit callbacks
- attachment_table.test.tsx: loading/empty states, row rendering, scope badges, delete admin gating
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude Code v2.1.69+ sends `custom: {defer_loading: true}` on tool
definitions. Anthropic's API accepts this field, but Bedrock rejects it
with "Extra inputs are not permitted", causing ~90% of requests to fail.
Strip the `custom` field from each tool in the request body before
sending to Bedrock, in both the Messages API and Chat API invoke paths.
Fixes#22847
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
* fix(proxy): readiness check returns 200 when database is unreachable
_db_health_readiness_check() catches health_check() exceptions but
never updates db_health_cache to "disconnected" and never re-raises.
The caller health_readiness() always returns 200 with "db": "connected"
hardcoded, regardless of actual DB state.
In Kubernetes, this means pods with dead database connections stay in
the Service endpoints and continue receiving traffic they cannot serve.
Changes:
- Set db_health_cache to "disconnected" and re-raise the exception on
health_check failure so health_readiness() returns 503
- Use actual db_health_status["status"] in the response instead of
hardcoding "db": "connected"
- Reduce cache TTL from 2 minutes to 15 seconds. The 2-minute window
is too wide for readiness probes (typically 10-15s intervals) and
means a pod can report healthy for up to 2 minutes after the DB dies
- Only serve cached results when status is "connected". The previous
condition (status != "unknown") would also cache "disconnected" for
2 minutes, delaying recovery detection after a DB comes back
* fix(proxy): add DB connection self-healing to readiness check
When the Prisma query engine's internal TCP connection pool holds dead
connections (caused by network blips, Cloud SQL proxy restarts, or
node-level issues), health_check() fails with httpx.ConnectError.
The engine never recovers on its own because nothing triggers a
disconnect/connect cycle to restart the subprocess with fresh
connections.
This leaves pods permanently failing readiness checks until they are
manually restarted, even after the underlying DB becomes reachable
again.
Add a reconnect attempt to _db_health_readiness_check() when
health_check() fails:
1. disconnect() - kills the query engine subprocess and closes all
connections (has built-in backoff retry: 3 tries, 10s max)
2. connect() - starts a new engine with fresh TCP connections (has
built-in backoff retry: 3 tries, 10s max)
3. health_check() - verifies the new connection works (has built-in
backoff retry: 3 tries, 10s max)
If reconnect succeeds, the pod immediately returns to service (200).
If it fails, the original exception is re-raised (503). Reconnect
attempts are rate-limited by probe frequency (~10-15s), so a
permanently unreachable DB gets one attempt per cycle with no retry
loops.
This uses the same disconnect/connect mechanism that
PrismaWrapper.recreate_prisma_client() uses for IAM token refresh,
and aligns with the community-documented pattern for Prisma connection
recovery in long-running processes (prisma/prisma#24718, #27024).
* Add poetry lock and modify test_health_endpoints
* Address allow_requests_on_db_unavailable regression
* Address comments
* resolve greptile issue
* Restore accidentally deleted UI HTML files
These were removed in an earlier commit but still exist on main.
Restoring to keep the PR diff clean.
* Guard reconnect with is_database_transport_error
Only attempt disconnect/connect/health_check cycle for transport-level
failures (unreachable DB, dropped connection). Data-layer errors like
UniqueViolationError indicate the DB is reachable, so reconnecting
would be pointless churn.
* Address greptile's comments
* Fix module alias after rebase and add adversarial test coverage
- Unify module alias to _health_endpoints_module after rebase conflict
- Add test for non-transport error with flag on (exercises is_database_transport_error guard)
- Add test for disconnect() failure during reconnect cycle
- Split non-transport error test into flag-off (re-raises) and flag-on (skips reconnect) variants
* Remove stale UI HTML files reintroduced during rebase
* fix: don't close HTTP/SDK clients on LLMClientCache eviction
Removing the _remove_key override that eagerly called aclose()/close()
on evicted clients. Evicted clients may still be held by in-flight
streaming requests; closing them causes:
RuntimeError: Cannot send a request, as the client has been closed.
This is a regression from commit fb72979432. Clients that are no longer
referenced will be garbage-collected naturally. Explicit shutdown cleanup
happens via close_litellm_async_clients().
Fixes production crashes after the 1-hour cache TTL expires.
* test: update LLMClientCache unit tests for no-close-on-eviction behavior
Flip the assertions: evicted clients must NOT be closed. Replace
test_remove_key_closes_async_client → test_remove_key_does_not_close_async_client
and equivalents for sync/eviction paths.
Add test_remove_key_removes_plain_values for non-client cache entries.
Remove test_background_tasks_cleaned_up_after_completion (no more _background_tasks).
Remove test_remove_key_no_event_loop variant that depended on old behavior.
* test: add e2e tests for OpenAI SDK client surviving cache eviction
Add two new e2e tests using real AsyncOpenAI clients:
- test_evicted_openai_sdk_client_stays_usable: verifies size-based eviction
doesn't close the client
- test_ttl_expired_openai_sdk_client_stays_usable: verifies TTL expiry
eviction doesn't close the client
Both tests sleep after eviction so any create_task()-based close would
have time to run, making the regression detectable.
Also expand the module docstring to explain why the sleep is required.
* docs(AGENTS.md): add rule — never close HTTP/SDK clients on cache eviction
* docs(CLAUDE.md): add HTTP client cache safety guideline
* Include user_email in new user creation within get_user_object
Enhance the get_user_object function to include user_email in the parameters when creating a new user. This change is accompanied by a new test to verify that user_email is correctly included during the upsert process.
* Improve error handling in test_get_user_object by logging exceptions
Updated the test_get_user_object_upsert_includes_user_email function to log exceptions when they occur, enhancing the visibility of potential issues during testing. This change helps in diagnosing failures related to the mock LiteLLM_UserTable.
* fix(passthrough): raise_for_status in _async_streaming to propagate Azure 429s
* address greptile review feedback (greploop iteration 1)
Guard data/json args when content is provided to avoid httpx ValueError
* address greptile review feedback (greploop iteration 2)
Use bare raise to preserve original traceback in _async_streaming exception handler
* address greptile review feedback (greploop iteration 3)
Close httpx streaming response on error to prevent connection pool exhaustion
* address greptile review feedback (greploop iteration 4)
Guard aclose() call to prevent masking original exception; add explicit test for content param forwarding
* address greptile review feedback (greploop iteration 5)
Pass content to sign_request so AWS body-hash signing is correct when content is the sole body source
* revert sign_request content change - request_data expects dict, not bytes
Bedrock's sign_request calls json.dumps(request_data) — passing content bytes
would TypeError. sign_request should only receive data/json (dict), not raw bytes.
All operational/diagnostic messages in WebSearchInterceptionLogger are now
debug-level to avoid flooding production logs while still remaining available
when verbose logging is enabled.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
PR #22890 used cast(str, ...) / cast(Optional[str], ...) for the return
statements; this PR's approach uses str() for explicit runtime coercion
(addressing Greptile's concern). Keep the str() version and drop the
now-unused cast import.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Per Sameerlite's review: warning-level logs trigger Slack alerts.
All 6 remaining .warning() calls were operational/fallback messages,
not actual errors. Changed to .info() to match the first fix at L510.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Documents exactly how every request and response field gets translated
when LiteLLM routes an Anthropic /v1/messages call through the OpenAI
Responses API path (for OpenAI/Azure targets). Covers messages content
block mapping, tools, tool_choice, thinking→reasoning, context_management,
and the reverse response translation. Wired into the /v1/messages sidebar.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- CreateBatchRequest.output_expires_after: drop Optional since total=False
already makes the key absent-or-present; Optional[T] incorrectly allowed
the key to exist with value None, which is incompatible with the OpenAI
SDK's OutputExpiresAfter | NotGiven expectation on batches.create()
- cost_tracking_settings._resolve_model_for_cost_lookup: replace implicit
object-to-str returns with explicit str() calls so the function is safe
even if the surrounding truthiness guards are later weakened
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>