litellm/tests
Curtis 725c0c158f
Prisma DB Failure Detection and Self-Healing (#21059)
* fix(proxy): readiness check returns 200 when database is unreachable

_db_health_readiness_check() catches health_check() exceptions but
never updates db_health_cache to "disconnected" and never re-raises.
The caller health_readiness() always returns 200 with "db": "connected"
hardcoded, regardless of actual DB state.

In Kubernetes, this means pods with dead database connections stay in
the Service endpoints and continue receiving traffic they cannot serve.

Changes:
- Set db_health_cache to "disconnected" and re-raise the exception on
  health_check failure so health_readiness() returns 503
- Use actual db_health_status["status"] in the response instead of
  hardcoding "db": "connected"
- Reduce cache TTL from 2 minutes to 15 seconds. The 2-minute window
  is too wide for readiness probes (typically 10-15s intervals) and
  means a pod can report healthy for up to 2 minutes after the DB dies
- Only serve cached results when status is "connected". The previous
  condition (status != "unknown") would also cache "disconnected" for
  2 minutes, delaying recovery detection after a DB comes back

* fix(proxy): add DB connection self-healing to readiness check

When the Prisma query engine's internal TCP connection pool holds dead
connections (caused by network blips, Cloud SQL proxy restarts, or
node-level issues), health_check() fails with httpx.ConnectError.
The engine never recovers on its own because nothing triggers a
disconnect/connect cycle to restart the subprocess with fresh
connections.

This leaves pods permanently failing readiness checks until they are
manually restarted, even after the underlying DB becomes reachable
again.

Add a reconnect attempt to _db_health_readiness_check() when
health_check() fails:
1. disconnect() - kills the query engine subprocess and closes all
   connections (has built-in backoff retry: 3 tries, 10s max)
2. connect() - starts a new engine with fresh TCP connections (has
   built-in backoff retry: 3 tries, 10s max)
3. health_check() - verifies the new connection works (has built-in
   backoff retry: 3 tries, 10s max)

If reconnect succeeds, the pod immediately returns to service (200).
If it fails, the original exception is re-raised (503). Reconnect
attempts are rate-limited by probe frequency (~10-15s), so a
permanently unreachable DB gets one attempt per cycle with no retry
loops.

This uses the same disconnect/connect mechanism that
PrismaWrapper.recreate_prisma_client() uses for IAM token refresh,
and aligns with the community-documented pattern for Prisma connection
recovery in long-running processes (prisma/prisma#24718, #27024).

* Add poetry lock and modify test_health_endpoints

* Address allow_requests_on_db_unavailable regression

* Address comments

* resolve greptile issue

* Restore accidentally deleted UI HTML files

These were removed in an earlier commit but still exist on main.
Restoring to keep the PR diff clean.

* Guard reconnect with is_database_transport_error

Only attempt disconnect/connect/health_check cycle for transport-level
failures (unreachable DB, dropped connection). Data-layer errors like
UniqueViolationError indicate the DB is reachable, so reconnecting
would be pointless churn.

* Address greptile's comments

* Fix module alias after rebase and add adversarial test coverage

- Unify module alias to _health_endpoints_module after rebase conflict
- Add test for non-transport error with flag on (exercises is_database_transport_error guard)
- Add test for disconnect() failure during reconnect cycle
- Split non-transport error test into flag-off (re-raises) and flag-on (skips reconnect) variants

* Remove stale UI HTML files reintroduced during rebase
2026-03-05 13:44:49 -08:00
..
agent_tests [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
audio_tests
basic_proxy_startup_tests
batches_tests Merge pull request #22625 from BerriAI/litellm_azure_ai_finetune 2026-03-03 19:42:17 +05:30
code_coverage_tests Merge pull request #22752 from BerriAI/litellm_search_api_add 2026-03-04 18:29:10 +05:30
documentation_tests fix failing tests 2026-02-21 15:48:26 -08:00
enterprise fix: support list of modes in Mode.default for tag-based guardrails 2026-03-04 01:50:28 +05:30
guardrails_tests fix: bump litellm-proxy-extras to 0.4.50 and fix 3 failing tests (#22417) 2026-02-28 10:20:03 -08:00
image_gen_tests
litellm Merge pull request #21233 from Chesars/feat/per-request-json-schema-validation 2026-03-03 15:29:54 -03:00
litellm_core_utils
litellm_utils_tests Fix test_perform_health_check_filters_by_model_id 2026-02-26 10:43:05 +05:30
litellm-proxy-extras
llm_responses_api_testing Fix anthropic responses 2026-02-20 17:30:42 -08:00
llm_translation fix(tools): gracefully repair truncated JSON in tool call arguments 2026-03-04 22:45:53 +01:00
load_tests
local_testing Add support for wildcards models for files api 2026-03-04 11:48:42 +05:30
logging_callback_tests Merge pull request #22180 from BerriAI/litellm_fix_vllm_test 2026-02-26 18:43:52 +05:30
mcp_tests Fix flaky MCP streaming test by properly mocking inner aresponses call 2026-03-04 11:09:24 -03:00
multi_instance_e2e_tests
ocr_tests
old_proxy_tests/tests
openai_endpoints_tests test(responses): add end-to-end test for responses API WebSocket mode 2026-03-02 17:24:39 +05:30
otel_tests
pass_through_tests Litellm stability fix v2 (#22452) 2026-02-28 15:29:45 -08:00
pass_through_unit_tests [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
proxy_admin_ui_tests security: fix critical/high CVEs in OS-level libs and NPM transitive 2026-02-24 19:40:09 +05:30
proxy_e2e_anthropic_messages_tests Fix: litellm/tests/llm_responses_api_testing/test_anthropic_responses_api.py 2026-02-20 17:30:53 -08:00
proxy_e2e_azure_batches_tests Add tenacity to e2e Azure batch CI and revert importorskip 2026-03-04 11:45:14 -03:00
proxy_security_tests
proxy_unit_tests Merge pull request #22372 from BerriAI/litellm_jwt_vkey_map 2026-03-05 06:24:49 +05:30
router_unit_tests [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
scim_tests
search_tests Add doc and tests for google search api 2026-03-04 13:55:43 +05:30
spend_tracking_tests
store_model_in_db_tests
test_litellm Prisma DB Failure Detection and Self-Healing (#21059) 2026-03-05 13:44:49 -08:00
unified_google_tests
vector_store_tests
windows_tests
__init__.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py
test_callbacks_on_proxy.py
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_openai_endpoints.py
test_organizations.py
test_otel_thread_leak.py
test_passthrough_endpoints.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_service_logger_otel.py
test_spend_logs.py
test_team_logging.py
test_team_members.py
test_team.py fix(test): skip 'projects' field in team update assertion (#21777) 2026-02-21 10:24:53 -08:00
test_users.py fix(tests): update deprecated Anthropic model in test_user_model_access (#21826) 2026-02-21 14:18:24 -08:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.