Skip to content

fix(llm_model_routing): sort OpenRouter providers by latency - #73

Open
mateusbellozupko wants to merge 1 commit into
evolution-foundation:mainfrom
mateusbellozupko:fix/openrouter-sort-by-latency
Open

mateusbellozupko wants to merge 1 commit into
evolution-foundation:mainfrom
mateusbellozupko:fix/openrouter-sort-by-latency

Conversation

@mateusbellozupko

@mateusbellozupko mateusbellozupko commented Sep 22, 2026 •

Copy link
Copy Markdown

Summary

  • OpenRouter's default provider routing balances price and speed for a given model, which can pick a cheaper-but-slower endpoint even when a faster one is available.
  • Added extra_body={"provider": {"sort": "latency"}} to the LiteLLM kwargs returned for provider="openrouter" agents. extra_body is LiteLLM's documented mechanism for forwarding OpenRouter-specific request fields straight through to the request body (see litellm/main.py's OpenrouterConfig handling — "we use openai 'extra_body' to pass openrouter specific params").
  • Verified against OpenRouter's official TypeScript SDK (OpenRouterTeam/ai-sdk-provider) that provider.sort is a real, typed field on the standard /chat/completions schema — as opposed to preferred_max_latency/preferred_min_throughput, which only exist on the separate Decisions/evaluation API and would have no effect on regular chat completions.
  • Found while investigating an incident where a single generation on an otherwise fast/reliable provider took ~104s and blew past the agent-run timeout mid-turn.

Test plan

  • pytest tests/unit/test_litellm_model_normalization.py — 11/11 passing (updated the existing parametrized assertions for the new extra_body key)

Summary by Sourcery

Route OpenRouter requests through the fastest available provider by default.

Bug Fixes:

  • Prioritize lower-latency OpenRouter providers for chat completions to avoid slow endpoint selection and agent-run timeouts.

Tests:

  • Update LiteLLM model normalization tests to verify the OpenRouter latency-routing configuration.

OpenRouter's default provider routing balances price and speed, which
can pick a cheaper-but-slower endpoint even when a faster one is
available for the same model. Added `extra_body={"provider": {"sort":
"latency"}}` to the LiteLLM kwargs for openrouter-routed agents —
LiteLLM's documented mechanism for forwarding OpenRouter-specific
request fields (see litellm/main.py's OpenrouterConfig handling).

Verified against OpenRouter's official TypeScript SDK
(OpenRouterTeam/ai-sdk-provider) that `provider.sort` is a real field
on the standard /chat/completions schema — as opposed to
`preferred_max_latency`, which only exists on the separate Decisions
API and would have no effect here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@sourcery-ai

sourcery-ai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Reviewer's guide (collapsed on small PRs)

Reviewer's Guide

OpenRouter model normalization now forwards provider.sort: latency via LiteLLM’s extra_body, biasing standard chat-completion routing toward faster endpoints while retaining existing API-base behavior. Unit-test expectations were updated to cover the new kwargs.

Sequence diagram for latency-based OpenRouter routing

sequenceDiagram
    participant Agent
    participant normalize_model_for_provider
    participant LiteLLM
    participant OpenRouter
    participant FastProvider

    Agent->>normalize_model_for_provider: normalize_model_for_provider(model, openrouter)
    normalize_model_for_provider-->>Agent: model, api_base, extra_body
    Agent->>LiteLLM: chat completion with extra_body
    LiteLLM->>OpenRouter: /chat/completions provider.sort=latency
    OpenRouter->>FastProvider: Route request to fastest endpoint
    FastProvider-->>OpenRouter: Generation response
    OpenRouter-->>LiteLLM: Completion response
    LiteLLM-->>Agent: Completion response
Loading

File-Level Changes

Change Details Files
Configure OpenRouter requests to prioritize the lowest-latency provider endpoint.
  • Preserve the OpenRouter API base URL configuration.
  • Forward provider.sort = latency through LiteLLM’s extra_body field.
  • Apply the routing options consistently, including when no model name is supplied.
src/utils/llm_model_routing.py
Update normalization tests to assert the new OpenRouter routing payload.
  • Extend expected kwargs with the nested latency-sorting extra_body value.
  • Retain empty kwargs assertions for non-OpenRouter providers.
tests/unit/test_litellm_model_normalization.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Approved.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant