fix(utils): allowed_openai_params must not forward unset params as None (#25777)

* feat(proxy): add NO_OPENAPI env var to disable /openapi.json endpoint (#25696)

* feat(proxy): add NO_OPENAPI env var to disable /openapi.json endpoint - Fixes #25538

* test(proxy): add tests for _get_openapi_url

---------

Co-authored-by: Progressive-engg <lov.kumari55@gmail.com>

* feat(prometheus): add api_provider label to spend metric (#25693)

* feat(prometheus): add api_provider label to spend metric

Add `api_provider` to `litellm_spend_metric` labels so users can
build Grafana dashboards that break down spend by cloud provider
(e.g. bedrock, anthropic, openai, azure, vertex_ai).

The `api_provider` label already exists in UserAPIKeyLabelValues and
is populated from `standard_logging_payload["custom_llm_provider"]`,
but was not included in the spend metric's label list.

* add api_provider to requests metric + add test

Address review feedback:
- Add api_provider to litellm_requests_metric too (same call-site as
  spend metric, keeps label sets in sync)
- Add test_api_provider_in_spend_and_requests_metrics following the
  existing pattern in test_prometheus_labels.py

* fix: ensure `litellm_metadata` is attached to `pre_call` guardrail to align with `post_call` guardrail (#25641)

* fix: ensure `litellm_metadata` is attached to pre_call to align with post_call

* refactor: remove unused BaseTranslation._ensure_litellm_metadata

* refactor: module level imports for ensure_litellm_metadata and CodeQL

* fix: update based off of Codex comment

* revert: undo usage of `_guardrail_litellm_metadata`

* feat: add pricing entry for openrouter/google/gemini-3.1-flash-lite-preview (#25610)

* fix(bedrock): skip synthetic tool injection for json_object with no schema (#25740)

When response_format={"type": "json_object"} is sent without a JSON
schema, _create_json_tool_call_for_response_format builds a tool with an
empty schema (properties: {}). The model follows the empty schema and
returns {} instead of the actual JSON the caller asked for.

This patch:
- Skips synthetic json_tool_call injection when no schema is provided.
  The model already returns JSON when the prompt asks for it.
- Fixes finish_reason: after _filter_json_mode_tools strips all
  synthetic tool calls, finish_reason stays "tool_calls" instead of
  "stop". Callers (like the OpenAI SDK) misinterpret this as a pending
  tool invocation.

json_schema requests with an explicit schema are unchanged.

Co-authored-by: Claude <noreply@anthropic.com>

* fix(utils): allowed_openai_params must not forward unset params as None

`_apply_openai_param_overrides` iterated `allowed_openai_params` and
unconditionally wrote `optional_params[param] = non_default_params.pop(param, None)`
for each entry. If the caller listed a param name but did not actually
send that param in the request, the pop returned `None` and `None` was
still written to `optional_params`. The openai SDK then rejected it as
a top-level kwarg:

    AsyncCompletions.create() got an unexpected keyword argument 'enable_thinking'

Reproducer (from #25697):

    allowed_openai_params = ["chat_template_kwargs", "enable_thinking"]
    body = {"chat_template_kwargs": {"enable_thinking": False}}

Here `enable_thinking` is only present nested inside
`chat_template_kwargs`, so the helper should forward
`chat_template_kwargs` and leave `enable_thinking` alone. Instead it
wrote `optional_params["enable_thinking"] = None`.

Fix: only forward a param if it was actually present in
`non_default_params`. Behavior is unchanged for the happy path (param
sent → still forwarded), and the explicit `None` leakage is gone.

Adds a regression test exercising the helper in isolation so the test
does not depend on any provider-specific `map_openai_params` plumbing.

Fixes #25697

---------

Co-authored-by: lovek629 <59618812+lovek629@users.noreply.github.com>
Co-authored-by: Progressive-engg <lov.kumari55@gmail.com>
Co-authored-by: Ori Kotek <ori.k@codium.ai>
Co-authored-by: Alexander Grattan <51346343+agrattan0820@users.noreply.github.com>
Co-authored-by: Mohana Siddhartha Chivukula <103447836+iamsiddhu3007@users.noreply.github.com>
Co-authored-by: Amiram Mizne <amiramm@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Tim Ren 2026-04-16 12:58:45 +08:00 committed by Sameer Kankute
parent 265a960472
commit dd4a41951f
No known key found for this signature in database
14 changed files with 478 additions and 12 deletions

View File

@ -1003,7 +1003,7 @@ class AmazonConverseConfig(BaseConfig):
description=description,
)
optional_params["outputConfig"] = output_config
else:
elif json_schema is not None:
# Fallback: translate to a synthetic tool call
# https://docs.anthropic.com/en/docs/build-with-claude/tool-use#json-mode
_tool = self._create_json_tool_call_for_response_format(
@ -1025,6 +1025,12 @@ class AmazonConverseConfig(BaseConfig):
)
if non_default_params.get("stream", False) is True:
optional_params["fake_stream"] = True
# else: response_format=json_object with no schema.
# Don't inject the synthetic json_tool_call tool here. When no
# schema is given, _create_json_tool_call_for_response_format
# produces an empty schema (properties: {}), and the model
# returns {} instead of the requested JSON. The model already
# returns JSON when the prompt asks for it.
optional_params["json_mode"] = True
return optional_params
@ -2030,6 +2036,12 @@ class AmazonConverseConfig(BaseConfig):
_message = Message(**chat_completion_message)
initial_finish_reason = map_finish_reason(completion_response["stopReason"])
# When json_mode filtered out all synthetic tool calls the response
# is plain content, not a pending tool invocation. Fix finish_reason
# so callers (e.g. OpenAI SDK) don't misinterpret it.
if json_mode and not filtered_tools and tools:
initial_finish_reason = "stop"
(
returned_message,
returned_finish_reason,

View File

@ -25179,6 +25179,58 @@
"supports_web_search": true,
"tpm": 800000
},
"openrouter/google/gemini-3.1-flash-lite-preview": {
"cache_read_input_token_cost": 2.5e-08,
"cache_read_input_token_cost_per_audio_token": 5e-08,
"input_cost_per_audio_token": 5e-07,
"input_cost_per_token": 2.5e-07,
"litellm_provider": "openrouter",
"max_audio_length_hours": 8.4,
"max_audio_per_prompt": 1,
"max_images_per_prompt": 3000,
"max_input_tokens": 1048576,
"max_output_tokens": 65536,
"max_pdf_size_mb": 30,
"max_tokens": 65536,
"max_video_length": 1,
"max_videos_per_prompt": 10,
"mode": "chat",
"output_cost_per_reasoning_token": 1.5e-06,
"output_cost_per_token": 1.5e-06,
"rpm": 2000,
"source": "https://ai.google.dev/pricing/gemini-3",
"supported_endpoints": [
"/v1/chat/completions",
"/v1/completions",
"/v1/batch"
],
"supported_modalities": [
"text",
"image",
"audio",
"video"
],
"supported_output_modalities": [
"text"
],
"supports_audio_input": true,
"supports_audio_output": false,
"supports_code_execution": true,
"supports_file_search": true,
"supports_function_calling": true,
"supports_parallel_function_calling": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_response_schema": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_url_context": true,
"supports_video_input": true,
"supports_vision": true,
"supports_web_search": true,
"tpm": 800000
},
"openrouter/google/gemini-3.1-pro-preview": {
"cache_read_input_token_cost": 2e-07,
"cache_read_input_token_cost_above_200k_tokens": 4e-07,

View File

@ -52,6 +52,20 @@ def _get_a2a_request_id(
endpoint_guardrail_translation_mappings = None
def _ensure_litellm_metadata(data: dict, user_api_key_dict: UserAPIKeyAuth) -> None:
"""Populate data['litellm_metadata'] from user_api_key_dict if absent."""
if "litellm_metadata" not in data:
from litellm.llms.base_llm.guardrail_translation.base_translation import (
BaseTranslation,
)
user_metadata = BaseTranslation.transform_user_api_key_dict_to_metadata(
user_api_key_dict
)
if user_metadata:
data["litellm_metadata"] = user_metadata
class UnifiedLLMGuardrails(CustomLogger):
def __init__(
self,
@ -120,6 +134,8 @@ class UnifiedLLMGuardrails(CustomLogger):
CallTypes(call_type)
]()
_ensure_litellm_metadata(data, user_api_key_dict)
data = await endpoint_translation.process_input_messages(
data=data,
guardrail_to_apply=guardrail_to_apply,
@ -177,6 +193,8 @@ class UnifiedLLMGuardrails(CustomLogger):
CallTypes(call_type)
]()
_ensure_litellm_metadata(data, user_api_key_dict)
return await endpoint_translation.process_input_messages(
data=data,
guardrail_to_apply=guardrail_to_apply,

View File

@ -496,6 +496,7 @@ from litellm.proxy.utils import (
_get_openapi_url,
_get_projected_spend_over_limit,
_get_redoc_url,
_get_openapi_url,
_is_projected_spend_over_limit,
_is_valid_team_configs,
get_custom_url,

View File

@ -5358,6 +5358,22 @@ def _get_docs_url() -> Optional[str]:
return "/"
def _get_openapi_url() -> Optional[str]:
"""
Get the OpenAPI JSON URL from the environment variables.
- If OPENAPI_URL is set, return it.
- If NO_OPENAPI is True, return None.
- Otherwise, default to "/openapi.json".
"""
if openapi_url := os.getenv("OPENAPI_URL"):
return openapi_url
if str_to_bool(os.getenv("NO_OPENAPI")) is True:
return None
return "/openapi.json"
def handle_exception_on_proxy(e: Exception) -> ProxyException:
"""

View File

@ -396,6 +396,7 @@ class PrometheusMetricLabels:
UserAPIKeyLabelNames.CLIENT_IP.value,
UserAPIKeyLabelNames.USER_AGENT.value,
UserAPIKeyLabelNames.MODEL_ID.value,
UserAPIKeyLabelNames.API_PROVIDER.value,
]
litellm_spend_metric = [
@ -410,6 +411,7 @@ class PrometheusMetricLabels:
UserAPIKeyLabelNames.CLIENT_IP.value,
UserAPIKeyLabelNames.USER_AGENT.value,
UserAPIKeyLabelNames.MODEL_ID.value,
UserAPIKeyLabelNames.API_PROVIDER.value,
]
litellm_input_tokens_metric = [

View File

@ -4811,11 +4811,21 @@ def _apply_openai_param_overrides(
If user passes in allowed_openai_params, apply them to optional_params
These params will get passed as is to the LLM API since the user opted in to passing them in the request
Only params the caller actually sent are forwarded. Previously this
function unconditionally wrote `None` for any allowed param missing from
the request, which then reached the provider SDK as a top-level kwarg it
did not recognize (e.g. openai SDK raising
`AsyncCompletions.create() got an unexpected keyword argument 'enable_thinking'`).
See https://github.com/BerriAI/litellm/issues/25697
"""
if allowed_openai_params:
for param in allowed_openai_params:
if param not in optional_params:
optional_params[param] = non_default_params.pop(param, None)
if param in optional_params:
continue
if param not in non_default_params:
continue
optional_params[param] = non_default_params.pop(param)
return optional_params

View File

@ -25164,6 +25164,58 @@
"supports_web_search": true,
"tpm": 800000
},
"openrouter/google/gemini-3.1-flash-lite-preview": {
"cache_read_input_token_cost": 2.5e-08,
"cache_read_input_token_cost_per_audio_token": 5e-08,
"input_cost_per_audio_token": 5e-07,
"input_cost_per_token": 2.5e-07,
"litellm_provider": "openrouter",
"max_audio_length_hours": 8.4,
"max_audio_per_prompt": 1,
"max_images_per_prompt": 3000,
"max_input_tokens": 1048576,
"max_output_tokens": 65536,
"max_pdf_size_mb": 30,
"max_tokens": 65536,
"max_video_length": 1,
"max_videos_per_prompt": 10,
"mode": "chat",
"output_cost_per_reasoning_token": 1.5e-06,
"output_cost_per_token": 1.5e-06,
"rpm": 2000,
"source": "https://ai.google.dev/pricing/gemini-3",
"supported_endpoints": [
"/v1/chat/completions",
"/v1/completions",
"/v1/batch"
],
"supported_modalities": [
"text",
"image",
"audio",
"video"
],
"supported_output_modalities": [
"text"
],
"supports_audio_input": true,
"supports_audio_output": false,
"supports_code_execution": true,
"supports_file_search": true,
"supports_function_calling": true,
"supports_parallel_function_calling": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
"supports_response_schema": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_url_context": true,
"supports_video_input": true,
"supports_vision": true,
"supports_web_search": true,
"tpm": 800000
},
"openrouter/google/gemini-3.1-pro-preview": {
"cache_read_input_token_cost": 2e-07,
"cache_read_input_token_cost_above_200k_tokens": 4e-07,

View File

@ -144,6 +144,45 @@ def test_get_optional_params_with_allowed_openai_params():
assert optional_params["reasoning_effort"] == reasoning_effort
def test_allowed_openai_params_does_not_forward_unset_params():
"""
Regression test for https://github.com/BerriAI/litellm/issues/25697
When a user lists a param in ``allowed_openai_params`` but does not
actually send that param in the request, litellm must not forward it
to the provider SDK as ``None``. The openai SDK rejects unknown
top-level kwargs with
``AsyncCompletions.create() got an unexpected keyword argument 'enable_thinking'``.
Reproduces the reported config where the user listed both
``chat_template_kwargs`` and ``enable_thinking`` in
``allowed_openai_params`` and only sent ``chat_template_kwargs``
(with ``enable_thinking`` nested inside it). Previously the loop
added ``optional_params["enable_thinking"] = None`` which then
crashed the openai client.
"""
from litellm.utils import _apply_openai_param_overrides
chat_template_kwargs = {"enable_thinking": False}
optional_params: dict = {}
non_default_params = {"chat_template_kwargs": chat_template_kwargs}
result = _apply_openai_param_overrides(
optional_params=optional_params,
non_default_params=non_default_params,
allowed_openai_params=["chat_template_kwargs", "enable_thinking"],
)
assert result["chat_template_kwargs"] == chat_template_kwargs
# enable_thinking was NOT sent as a top-level param — it must not be
# forwarded to the provider SDK (openai AsyncCompletions.create would
# reject an unknown kwarg, even if its value is None).
assert "enable_thinking" not in result
# And the only entry actually moved out of non_default_params is
# the one the caller sent.
assert "chat_template_kwargs" not in non_default_params
def test_bedrock_optional_params_embeddings():
litellm.drop_params = True
optional_params = get_optional_params_embeddings(

View File

@ -10,7 +10,7 @@ import pytest
from fastapi import Request
from starlette.datastructures import State
from litellm.proxy.utils import _get_docs_url, _get_redoc_url
from litellm.proxy.utils import _get_docs_url, _get_openapi_url, _get_redoc_url
sys.path.insert(
0, os.path.abspath("../..")
@ -735,6 +735,30 @@ def test_get_docs_url(env_vars, expected_url):
result = _get_docs_url()
assert result == expected_url
@pytest.mark.parametrize(
"env_vars, expected_url",
[
({}, "/openapi.json"), # default case
({"OPENAPI_URL": "/custom-openapi.json"}, "/custom-openapi.json"), # custom URL
(
{"OPENAPI_URL": "https://example.com/openapi.json"},
"https://example.com/openapi.json",
), # full URL
({"NO_OPENAPI": "True"}, None), # openapi disabled
],
)
def test_get_openapi_url(env_vars, expected_url):
# Clear relevant environment variables
for key in ["OPENAPI_URL", "NO_OPENAPI"]:
os.environ.pop(key, None)
# Set test environment variables
for key, value in env_vars.items():
os.environ[key] = value
result = _get_openapi_url()
assert result == expected_url
@pytest.mark.parametrize(
"request_tags, tags_to_add, expected_tags",

View File

@ -69,6 +69,21 @@ def test_model_id_in_required_metrics():
print(f"{metric_name} contains model_id label")
def test_api_provider_in_spend_and_requests_metrics():
"""
Test that api_provider label is present in spend and requests metrics
so users can build spend-by-provider and request-count-by-provider dashboards.
"""
api_provider_label = UserAPIKeyLabelNames.API_PROVIDER.value
for metric_name in ["litellm_spend_metric", "litellm_requests_metric"]:
labels = PrometheusMetricLabels.get_labels(metric_name)
assert (
api_provider_label in labels
), f"Metric {metric_name} should contain api_provider label"
print(f"{metric_name} contains api_provider label")
def test_user_email_label_exists():
"""Test that the USER_EMAIL label is properly defined"""
assert UserAPIKeyLabelNames.USER_EMAIL.value == "user_email"

View File

@ -3140,9 +3140,14 @@ def test_add_additional_properties_definitions():
assert result["definitions"]["Item"]["properties"]["details"]["additionalProperties"] is False
def test_json_object_no_schema_falls_back_to_tool_call():
"""response_format: {type: json_object} with no schema should use tool-call fallback,
even for models that support native structured outputs."""
def test_json_object_no_schema_skips_tool_injection():
"""response_format: {type: json_object} with no schema should NOT inject
the synthetic json_tool_call tool.
When no schema is given, _create_json_tool_call_for_response_format builds
a tool with an empty schema (properties: {}). The model follows the schema
and returns {} instead of the requested JSON. Skipping tool injection lets
the model respond naturally with the JSON the caller asked for."""
old_env = os.environ.get("LITELLM_LOCAL_MODEL_COST_MAP")
old_cost = litellm.model_cost
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
@ -3162,8 +3167,9 @@ def test_json_object_no_schema_falls_back_to_tool_call():
# Should NOT use native outputConfig (no schema provided)
assert "outputConfig" not in result
# Should use tool-call fallback
assert "tools" in result
# Should NOT inject tools - empty schema causes model to return {}
assert "tools" not in result
assert "tool_choice" not in result
assert result["json_mode"] is True
finally:
litellm.model_cost = old_cost
@ -3655,3 +3661,140 @@ def test_cache_control_injection_tool_config_not_added_without_injection_point()
tools = result["toolConfig"]["tools"]
# No cachePoint should be appended
assert all("cachePoint" not in tool for tool in tools)
def test_translate_response_format_json_schema_still_injects_tool():
"""
response_format with an explicit json_schema should still use the
synthetic tool call approach (for models that don't support native
structured outputs).
"""
config = AmazonConverseConfig()
response_format = {
"type": "json_schema",
"json_schema": {
"name": "FactResult",
"schema": {
"type": "object",
"properties": {
"facts": {
"type": "array",
"items": {"type": "string"},
},
},
"required": ["facts"],
},
},
}
optional_params: dict = {}
result = config._translate_response_format_param(
value=response_format,
model="anthropic.claude-3-haiku-20240307-v1:0",
optional_params=optional_params,
non_default_params={"response_format": response_format},
is_thinking_enabled=False,
)
assert result["json_mode"] is True
assert "tools" in result
assert "tool_choice" in result
def test_transform_response_finish_reason_stop_when_json_mode_filters_all_tools():
"""
When json_mode is True and _filter_json_mode_tools strips all synthetic
tool calls, finish_reason should be "stop", not "tool_calls".
Bedrock returns stopReason="tool_use" for json_tool_call responses.
After filtering, the response is plain JSON content and should not look
like a pending tool invocation to callers.
"""
from litellm.llms.bedrock.chat.converse_transformation import AmazonConverseConfig
from litellm.types.utils import ModelResponse
response_json = {
"metrics": {"latencyMs": 100},
"output": {
"message": {
"role": "assistant",
"content": [
{
"toolUse": {
"toolUseId": "tooluse_001",
"name": "json_tool_call",
"input": {
"facts": ["Bob is a software engineer"],
},
}
}
],
}
},
"stopReason": "tool_use",
"usage": {
"inputTokens": 50,
"outputTokens": 20,
"totalTokens": 70,
"cacheReadInputTokenCount": 0,
"cacheReadInputTokens": 0,
"cacheWriteInputTokenCount": 0,
"cacheWriteInputTokens": 0,
},
}
class MockResponse:
def json(self):
return response_json
@property
def text(self):
return json.dumps(response_json)
config = AmazonConverseConfig()
model_response = ModelResponse()
# Simulate what happens when json_tool_call was injected for a
# json_schema request: optional_params has the synthetic tool
optional_params = {
"json_mode": True,
"tools": [
{
"type": "function",
"function": {
"name": "json_tool_call",
"parameters": {
"type": "object",
"additionalProperties": True,
"properties": {},
},
},
}
],
}
result = config._transform_response(
model="bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
response=MockResponse(),
model_response=model_response,
stream=False,
logging_obj=None,
optional_params=optional_params,
api_key=None,
data=None,
messages=[],
encoding=None,
)
# Content should have the JSON from the tool call arguments
content = result.choices[0].message.content
assert content is not None
parsed = json.loads(content)
assert parsed["facts"] == ["Bob is a software engineer"]
# No tool_calls on the message
assert result.choices[0].message.tool_calls is None
# finish_reason must be "stop", not "tool_calls"
assert result.choices[0].finish_reason == "stop"

View File

@ -4,6 +4,7 @@ import pytest
from fastapi import HTTPException
from litellm.integrations.custom_guardrail import ModifyResponseException
from litellm.proxy._types import UserAPIKeyAuth
from litellm.proxy.guardrails.guardrail_hooks.grayswan.grayswan import (
GraySwanGuardrail,
GraySwanGuardrailAPIError,
@ -435,7 +436,10 @@ def test_format_violation_message() -> None:
assert "Gray Swan Cygnal Guardrail" in message
assert "the input query has a violation score of 0.85" in message
assert "violating the rule(s): 1, 3, 5" in message
assert "Mutation effort to make the harmful intention disguised was DETECTED" in message
assert (
"Mutation effort to make the harmful intention disguised was DETECTED"
in message
)
# IPI should not be in message since it's False
assert "Indirect Prompt Injection was DETECTED" not in message
@ -446,4 +450,58 @@ def test_format_violation_message() -> None:
assert "Gray Swan Cygnal Guardrail" in message
assert "the model response has a violation score of 0.85" in message
assert "violating the rule(s): 1, 3, 5" in message
assert "Mutation effort to make the harmful intention disguised was DETECTED" in message
assert (
"Mutation effort to make the harmful intention disguised was DETECTED"
in message
)
def test_prepare_payload_includes_litellm_metadata(
grayswan_guardrail: GraySwanGuardrail,
) -> None:
"""Verify _prepare_payload forwards litellm_metadata from request_data."""
messages = [{"role": "user", "content": "hello"}]
request_data = {
"litellm_metadata": {
"user_api_key_user_id": "user-123",
"user_api_key_team_id": "team-456",
"user_api_key_spend": 0,
}
}
payload = grayswan_guardrail._prepare_payload(messages, {}, request_data)
assert payload is not None
assert "litellm_metadata" in payload
assert payload["litellm_metadata"]["user_api_key_user_id"] == "user-123"
assert payload["litellm_metadata"]["user_api_key_team_id"] == "team-456"
def test_ensure_litellm_metadata_populates_from_user_api_key_dict() -> None:
"""Verify _ensure_litellm_metadata populates litellm_metadata."""
from litellm.proxy.guardrails.guardrail_hooks.unified_guardrail.unified_guardrail import (
_ensure_litellm_metadata,
)
user_auth = UserAPIKeyAuth(user_id="u1", team_id="t1", api_key="sk-test-hashed")
data: dict = {}
_ensure_litellm_metadata(data, user_auth)
assert "litellm_metadata" in data
assert data["litellm_metadata"]["user_api_key_user_id"] == "u1"
assert data["litellm_metadata"]["user_api_key_team_id"] == "t1"
def test_ensure_litellm_metadata_noop_when_already_present() -> None:
"""Verify _ensure_litellm_metadata does not overwrite existing litellm_metadata."""
from litellm.proxy.guardrails.guardrail_hooks.unified_guardrail.unified_guardrail import (
_ensure_litellm_metadata,
)
user_auth = UserAPIKeyAuth(user_id="should-not-appear")
data: dict = {"litellm_metadata": {"existing": "value"}}
_ensure_litellm_metadata(data, user_auth)
assert data["litellm_metadata"] == {"existing": "value"}

View File

@ -1876,7 +1876,7 @@ def test_gemini_without_cache_tokens_details():
"promptTokensDetails": [
{"modality": "TEXT", "tokenCount": 6},
{"modality": "IMAGE", "tokenCount": 258},
]
],
# No cacheTokensDetails
}
}
@ -2014,3 +2014,27 @@ def test_additional_costs_only_for_azure_ai():
completion_tokens=50,
)
assert result is None, "Vertex AI should have no additional costs"
def test_openrouter_gemini_3_1_flash_lite_preview_pricing():
"""
Test that openrouter/google/gemini-3.1-flash-lite-preview has a pricing entry.
Regression test for https://github.com/BerriAI/litellm/issues/25604
The model exists and is callable via OpenRouter, but was missing from
model_prices_and_context_window.json when other Gemini 3.x variants were present.
This caused ValueError: This model isn't mapped yet during router pre-call checks.
"""
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
model_name = "openrouter/google/gemini-3.1-flash-lite-preview"
model_info = litellm.model_cost.get(model_name)
assert model_info is not None, f"Missing model pricing entry: {model_name}"
assert model_info["litellm_provider"] == "openrouter"
assert model_info["input_cost_per_token"] == 2.5e-07
assert model_info["output_cost_per_token"] == 1.5e-06
assert model_info["max_input_tokens"] == 1048576
assert model_info["max_output_tokens"] == 65536