Repository navigation
fix(llm_model_routing): sort OpenRouter providers by latency - #73
Open
mateusbellozupko wants to merge 1 commit into
Open
mateusbellozupko wants to merge 1 commit into
mateusbellozupko wants to merge 1 commit into
Conversation
OpenRouter's default provider routing balances price and speed, which
can pick a cheaper-but-slower endpoint even when a faster one is
available for the same model. Added `extra_body={"provider": {"sort":
"latency"}}` to the LiteLLM kwargs for openrouter-routed agents —
LiteLLM's documented mechanism for forwarding OpenRouter-specific
request fields (see litellm/main.py's OpenrouterConfig handling).
Verified against OpenRouter's official TypeScript SDK
(OpenRouterTeam/ai-sdk-provider) that `provider.sort` is a real field
on the standard /chat/completions schema — as opposed to
`preferred_max_latency`, which only exists on the separate Decisions
API and would have no effect here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reviewer's guide (collapsed on small PRs)Reviewer's GuideOpenRouter model normalization now forwards Sequence diagram for latency-based OpenRouter routingsequenceDiagram
participant Agent
participant normalize_model_for_provider
participant LiteLLM
participant OpenRouter
participant FastProvider
Agent->>normalize_model_for_provider: normalize_model_for_provider(model, openrouter)
normalize_model_for_provider-->>Agent: model, api_base, extra_body
Agent->>LiteLLM: chat completion with extra_body
LiteLLM->>OpenRouter: /chat/completions provider.sort=latency
OpenRouter->>FastProvider: Route request to fastest endpoint
FastProvider-->>OpenRouter: Generation response
OpenRouter-->>LiteLLM: Completion response
LiteLLM-->>Agent: Completion response
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
extra_body={"provider": {"sort": "latency"}}to the LiteLLM kwargs returned forprovider="openrouter"agents.extra_bodyis LiteLLM's documented mechanism for forwarding OpenRouter-specific request fields straight through to the request body (seelitellm/main.py'sOpenrouterConfighandling — "we use openai 'extra_body' to pass openrouter specific params").OpenRouterTeam/ai-sdk-provider) thatprovider.sortis a real, typed field on the standard/chat/completionsschema — as opposed topreferred_max_latency/preferred_min_throughput, which only exist on the separate Decisions/evaluation API and would have no effect on regular chat completions.Test plan
pytest tests/unit/test_litellm_model_normalization.py— 11/11 passing (updated the existing parametrized assertions for the newextra_bodykey)Summary by Sourcery
Route OpenRouter requests through the fastest available provider by default.
Bug Fixes:
Tests: