litellm/tests/test_litellm
Curtis 725c0c158f
Prisma DB Failure Detection and Self-Healing (#21059)
* fix(proxy): readiness check returns 200 when database is unreachable

_db_health_readiness_check() catches health_check() exceptions but
never updates db_health_cache to "disconnected" and never re-raises.
The caller health_readiness() always returns 200 with "db": "connected"
hardcoded, regardless of actual DB state.

In Kubernetes, this means pods with dead database connections stay in
the Service endpoints and continue receiving traffic they cannot serve.

Changes:
- Set db_health_cache to "disconnected" and re-raise the exception on
  health_check failure so health_readiness() returns 503
- Use actual db_health_status["status"] in the response instead of
  hardcoding "db": "connected"
- Reduce cache TTL from 2 minutes to 15 seconds. The 2-minute window
  is too wide for readiness probes (typically 10-15s intervals) and
  means a pod can report healthy for up to 2 minutes after the DB dies
- Only serve cached results when status is "connected". The previous
  condition (status != "unknown") would also cache "disconnected" for
  2 minutes, delaying recovery detection after a DB comes back

* fix(proxy): add DB connection self-healing to readiness check

When the Prisma query engine's internal TCP connection pool holds dead
connections (caused by network blips, Cloud SQL proxy restarts, or
node-level issues), health_check() fails with httpx.ConnectError.
The engine never recovers on its own because nothing triggers a
disconnect/connect cycle to restart the subprocess with fresh
connections.

This leaves pods permanently failing readiness checks until they are
manually restarted, even after the underlying DB becomes reachable
again.

Add a reconnect attempt to _db_health_readiness_check() when
health_check() fails:
1. disconnect() - kills the query engine subprocess and closes all
   connections (has built-in backoff retry: 3 tries, 10s max)
2. connect() - starts a new engine with fresh TCP connections (has
   built-in backoff retry: 3 tries, 10s max)
3. health_check() - verifies the new connection works (has built-in
   backoff retry: 3 tries, 10s max)

If reconnect succeeds, the pod immediately returns to service (200).
If it fails, the original exception is re-raised (503). Reconnect
attempts are rate-limited by probe frequency (~10-15s), so a
permanently unreachable DB gets one attempt per cycle with no retry
loops.

This uses the same disconnect/connect mechanism that
PrismaWrapper.recreate_prisma_client() uses for IAM token refresh,
and aligns with the community-documented pattern for Prisma connection
recovery in long-running processes (prisma/prisma#24718, #27024).

* Add poetry lock and modify test_health_endpoints

* Address allow_requests_on_db_unavailable regression

* Address comments

* resolve greptile issue

* Restore accidentally deleted UI HTML files

These were removed in an earlier commit but still exist on main.
Restoring to keep the PR diff clean.

* Guard reconnect with is_database_transport_error

Only attempt disconnect/connect/health_check cycle for transport-level
failures (unreachable DB, dropped connection). Data-layer errors like
UniqueViolationError indicate the DB is reachable, so reconnecting
would be pointless churn.

* Address greptile's comments

* Fix module alias after rebase and add adversarial test coverage

- Unify module alias to _health_endpoints_module after rebase conflict
- Add test for non-transport error with flag on (exercises is_database_transport_error guard)
- Add test for disconnect() failure during reconnect cycle
- Split non-transport error test into flag-off (re-raises) and flag-on (skips reconnect) variants

* Remove stale UI HTML files reintroduced during rebase
2026-03-05 13:44:49 -08:00
..
a2a_protocol
anthropic_interface/exceptions
caching fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22925) 2026-03-05 12:00:38 -08:00
completion_extras test(streaming): add comprehensive parallel tool call integration test 2026-03-04 10:17:20 +00:00
containers fix: add missing OpenAI chat completion params to OPENAI_CHAT_COMPLETION_PARAMS (#21360) 2026-02-16 20:31:21 -08:00
enterprise Fixes based on greptile reviews 2026-02-18 12:19:11 +05:30
expected_responses_api_request [Feat] Adds support for server-side compaction on the OpenAI Responses API context_management (#21058) 2026-02-12 10:00:30 -08:00
experimental_mcp_client
google_genai
images Merge pull request #22307 from Chesars/fix/22244-image-edit-custom-pricing 2026-02-27 16:38:34 -03:00
integrations fix: replace sk-fake with safe test key to avoid secret scanner 2026-03-05 18:29:28 +01:00
interactions fix(test): Update status enum values to match Google Interactions OpenAPI spec (#22061) 2026-02-24 20:26:11 -08:00
litellm_core_utils fix(streaming): prevent Vertex AI Claude content truncation when finish_reason races content 2026-03-04 22:44:48 +01:00
llms feat(anthropic): support top-level cache_control for automatic prompt caching (#22442) 2026-03-05 08:34:56 -08:00
ocr Enable local file support for OCR (#22133) 2026-02-27 10:50:02 -08:00
passthrough fix(passthrough): propagate Azure 429/5xx errors in async streaming instead of silent HTTP 200 (#22913) 2026-03-05 10:12:43 -08:00
proxy Prisma DB Failure Detection and Self-Healing (#21059) 2026-03-05 13:44:49 -08:00
responses Add tests for resp 2026-03-04 18:27:06 +05:30
router_strategy fix: complexity_router crashes on list-format message content (OpenAI multi-part messages) (#22761) 2026-03-04 16:18:49 -08:00
router_utils Fix encrypted content streaming affinity issue 2026-03-03 18:37:22 +05:30
secret_managers fix(tests): isolate flaky files endpoint tests from global proxy state (#21788) 2026-02-21 11:20:32 -08:00
test_router fix: use atomic increment-first pattern for model RPM rate limiting 2026-02-24 09:55:07 -03:00
types fix: add video_tokens to expected completion_tokens_details in test 2026-03-03 19:46:20 -03:00
vector_stores
__init__.py
conftest.py fix(tests): restore disable_aiohttp_transport and force_ipv4 in isolate_litellm_state 2026-02-17 21:18:49 -03:00
log.txt
readme.md
test_a2a_registry_lookup.py
test_acompletion_session_reuse_e2e.py
test_add_deployment_no_master_key.py
test_aembedding_session_reuse_e2e.py
test_anthropic_beta_headers_filtering.py Make tests run with local beta header mapping json 2026-02-13 22:31:42 +05:30
test_azure_video_router.py
test_claude_haiku_4_5_config.py
test_claude_opus_4_6_config.py Fix au.anthropic.claude opus 4 6 v1 (#20731) 2026-02-16 14:15:37 -08:00
test_constants.py added configurable env for mcp timeouts (#22287) 2026-03-02 13:13:41 -08:00
test_container_router.py
test_cost_calculation_log_level.py fix(tests): use record.getMessage() instead of record.message for LogRecord 2026-02-18 11:46:32 -03:00
test_cost_calculator.py perf: optimize completion_cost() — eliminate enum overhead, reduce function call indirection 2026-02-21 12:14:55 -08:00
test_deepseek_model_metadata.py
test_eager_tiktoken_load.py
test_exception_exports.py
test_exception_header_preservation.py Update test to righ place 2026-02-26 13:26:51 -08:00
test_exception_mapping_request_attribute.py
test_filter_out_litellm_params.py
test_get_blog_posts.py fix(ollama): thread api_base to get_model_info + graceful fallback (#21970) 2026-02-23 21:00:37 -08:00
test_gpt_image_cost_calculator.py
test_groq_streaming_encoding.py
test_lazy_imports.py
test_logging.py
test_lowest_latency_zero_tokens.py
test_main.py Fix : test_video_content_handler_uses_get_for_openai 2026-02-17 20:06:08 +05:30
test_model_param_helper.py
test_model_response_normalization.py fix(types): remove StreamingChoices from ModelResponse, use ModelResponseStream 2026-02-20 17:47:42 -03:00
test_nested_drop_params.py
test_project_tags_pydantic.py fix: req changes 2026-02-27 13:33:34 +05:30
test_redis.py
test_register_model_custom_pricing.py test: fix misleading precedence test per review feedback 2026-03-02 08:31:33 +00:00
test_responses_api_bridge_non_stream.py
test_responses_id_security.py Fix responses ID security test for new request_cache parameter 2026-03-04 11:29:51 -03:00
test_router_google_genai.py
test_router_model_cost_isolation.py
test_router_per_deployment_num_retries.py
test_router_redis_init.py
test_router_silent_experiment.py
test_router.py fix: add sync streaming fallback + fix 429 for all streaming paths (#22375) 2026-02-28 15:55:05 -08:00
test_service_logger.py fix(proxy): fix master key rotation Prisma validation errors (#21330) 2026-02-16 15:13:05 -08:00
test_shared_session_integration.py
test_ssl_verify_unit.py BUMP Enterprise PIP 2026-02-14 13:40:48 -08:00
test_streaming_connection_cleanup.py fix: add debug logging to stream cleanup, improve tests 2026-02-14 17:31:39 -08:00
test_system_message_format_bug.py
test_utils.py fix(test): add 'realtime' to model mode enum in schema validation 2026-03-05 06:41:51 -03:00
test_uuid_helper.py
test_video_generation.py fix(ollama): thread api_base to get_model_info + graceful fallback (#21970) 2026-02-23 21:00:37 -08:00
test_xai_responses_auto_routing.py

Testing for litellm/

This directory 1:1 maps the the litellm/ directory, and can only contain mocked tests.

The point of this is to:

  1. Increase test coverage of litellm/
  2. Make it easy for contributors to add tests for the litellm/ package and easily run tests without needing LLM API keys.

File name conventions

  • litellm/proxy/test_caching_routes.py maps to litellm/proxy/caching_routes.py
  • test_<filename>.py maps to litellm/<filename>.py