litellm

History

Sameer Kankute 477b63c5ea fix(caching): replay openai/responses bridge cache hits as chat streams (#28158 ) * fix(caching): replay openai/responses bridge cache hits as chat streams When chat completions route through openai/responses, cached ModelResponse payloads under aresponses keys were deserialized as ResponsesAPIResponse (500) or re-translated as responses events (empty streaming deltas). Deserialize chat-shaped cache entries as acompletion and bypass the responses stream iterator for cached CustomStreamWrapper replay. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(caching): map responses bridge call_type for sync vs async stream replay Co-authored-by: Yassin Kortam <yassin@berri.ai> * fix: handle ModelResponse cache return in responses bridge and drop dead acompletion check Co-authored-by: Yassin Kortam <yassin@berri.ai> * fix(caching): detect chat cache hits via object field before choices fallback Prefer chat.completion object type over the broad choices-key heuristic so Responses API cached payloads are not misclassified if their schema changes. Co-authored-by: Cursor <cursoragent@cursor.com> * test(caching): cover responses bridge cache-hit paths in CI-tracked test suite The new bridge cache replay logic in caching_handler.py and the preformatted-stream guard in litellm_responses_transformation/handler.py were exercised only by tests under tests/local_testing/, which the responses-caching-types and misc shards do not run. Codecov flagged the patch as 29.72% covered. Add equivalent unit tests under tests/test_litellm/ so the responses, caching, types, and misc shards execute them and ship their coverage data to Codecov: - _is_chat_completion_cached_dict happy/sad paths - aresponses streaming bridge cache hit -> CustomStreamWrapper - responses non-streaming bridge cache hit -> ModelResponse - legacy ResponsesAPIResponse stream + non-stream replay - _is_preformatted_cached_chat_stream true/false - completion/acompletion early return on cached ModelResponse - completion/acompletion skip rewrap on preformatted cached stream * fix: add negative guard on object field in _is_chat_completion_cached_dict Co-authored-by: Yassin Kortam <yassin@berri.ai> * fix(vcr): treat corrupt cassette payloads as cache miss * test: bump EOL'd NVIDIA rerank and OpenAI realtime models in CI The NVIDIA hosted rerank endpoint for nvidia/llama-3_2-nv-rerankqa-1b-v2 reached end-of-life on 2026-05-18 and now returns HTTP 410 Gone, breaking TestNvidiaNim::test_basic_rerank. Switch to nvidia/nv-rerankqa-mistral-4b-v3, which is still hosted on the NVIDIA API catalog and is already listed in model_prices_and_context_window.json. OpenAI also retired the gpt-4o-realtime-preview-2024-12-17 model used by test_realtime_guardrails_openai (now returns model_not_found). Switch the realtime test URL to the GA gpt-realtime alias. Unrelated to the responses-bridge cache fix in this PR, but committing here to unblock CI per maintainer guidance. Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com> * test(realtime): switch retired gpt-4o-realtime-preview to gpt-realtime OpenAI removed gpt-4o-realtime-preview and all its date snapshots on 2026-05-18 (every variant now returns model_not_found), breaking the live-WebSocket OpenAI realtime tests in CI: - test_openai_realtime_direct_call_no_intent - test_openai_realtime_direct_call_with_intent - TestOpenAIRealtime.test_realtime_connection - TestOpenAIRealtime.test_realtime_with_query_params Point each of those to the current GA alias gpt-realtime (verified live). Pure unit/mock tests that just assert the string value (e.g. in test_realtime_query_params_construction and the test_realtime_query_params_use_normalized_model_name mock) are left alone since they do not depend on model availability. Also relax the AI-response assertion in test_text_message_blocked_by_guardrail_no_ai_response: gpt-realtime occasionally produces a polite refusal ("I'm sorry, but I can't say that") when the cancel arrives after the model has already started generating, which is the expected outcome (no real AI content) but does not contain the words 'blocked' or 'guardrail'. The primary guardrail behaviour (guardrail_violation error event + transcript_delta block message) is still asserted unchanged. Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com> * test(nvidia_nim): mock rerank live API instead of hitting EOL'd endpoint NVIDIA reached end-of-life for the hosted nvidia/llama-3.2-nv-rerankqa-1b-v2 rerank API on 2026-05-18 (returns HTTP 410 Gone), and the proposed replacement nv-rerankqa-mistral-4b-v3 returns HTTP 404 for the CI account, breaking TestNvidiaNim::test_basic_rerank. Override test_basic_rerank to mock the HTTP transport (same pattern as test_nvidia_nim_rerank_ranking_endpoint above) so the request/response transformation and cost calculation stay covered without depending on NVIDIA's hosted catalog rotation. The model identifier reverts to the original llama-3.2-nv-rerankqa-1b-v2 since the request never leaves the test process. --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Yassin Kortam <yassin@berri.ai> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>		2026-05-18 16:27:06 -07:00
..
test_azure_blob_cache.py	style: run black formatter on files from main merge	2026-04-17 13:02:59 -07:00
test_caching_handler.py	fix(caching): replay openai/responses bridge cache hits as chat streams (#28158 )	2026-05-18 16:27:06 -07:00
test_dual_cache.py	update tests	2026-04-21 23:30:46 +00:00
test_gcs_cache.py	style: run black formatter on files from main merge	2026-04-17 13:02:59 -07:00
test_in_memory_cache.py	fix: prune expired in-memory cache heap entries (#25664 )	2026-04-14 23:37:49 +05:30
test_llm_caching_handler.py	fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22925 )	2026-03-05 12:00:38 -08:00
test_llm_client_cache_e2e.py	fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22925 )	2026-03-05 12:00:38 -08:00
test_qdrant_semantic_cache.py	chore(caching): remove allow_legacy_unscoped_cache_hits opt-in	2026-05-04 22:16:30 +00:00
test_redis_cache.py	Refresh Redis TTL on counter writes and skip stale in-memory on Redis miss	2026-04-30 17:50:58 -07:00
test_redis_cluster_cache.py	fix(caching): check REDIS_CLUSTER_NODES env var in Cache and Router class selection (#22790 )	2026-03-06 17:31:30 -08:00
test_redis_connection_pool.py	style: run black formatter on files from main merge	2026-04-17 13:02:59 -07:00
test_redis_semantic_cache.py	chore(caching): remove allow_legacy_unscoped_cache_hits opt-in	2026-05-04 22:16:30 +00:00
test_s3_cache.py	style: run black formatter on files from main merge	2026-04-17 13:02:59 -07:00