Skip to content
American owned and operated
Nightshift

Semantic Router

Every request.
The right model.
Every time.

Some questions need your most capable model. Most do not. The router reads each one, sends it where it belongs, and shows you what it cost.

The problem

Your hardest 10% of business requests shouldn’t determine the cost of the other 90% on your model.

Most business requests don’t require frontier-level AI. Nightshift provides detailed use, cost and performance reporting by category and routes each one to the most efficient model, data source, or workflow— using premium intelligence only when the value justifies it.

  • Where is the value?

    Can we connect each AI use case to revenue, productivity, cost reduction, or risk?

  • Is this the right model for the work?

    Different tasks require different levels of intelligence, speed, specialization, and data access.

  • Can we trust it in production?

    Are the answers accurate, the service reliable, and the decisions observable, with human oversight where it matters?

How it works

The request determines the route.

Nightshift SR reads every prompt, checks it against your rules, and sends it to the model that should answer it.

Your environment

Summarize this week’s support tickets.

Simple, low risk.

Semantic Router
  • 01

    Classify keywords

  • 02

    Compare against rules

  • 03

    Reveal business context

  • 04

    Determine complexity

  • 05

    Decide model route

Your mixture of models

Closed source

  • OpenAI
  • Anthropic
  • Gemini

Open source

  • DeepSeek
  • Llama
  • Mistral

Balancing

It chooses a model on every request.

Not once, when your systems were built. Every time, and from whatever is healthy at that moment. The router has no favorite, and the provider you already use is one of the places it can send a request.

  • Scores what matters

    Quality, price and live speed, weighed for the prompt in front of it.

  • Drops what is struggling

    A model that is throttled or slow stops receiving traffic until it recovers.

  • Spends your commitment first

    Traffic goes to the capacity you already pay for before anything metered by the token.

  • Nobody notices

    A provider having a bad afternoon becomes a different model answering, rather than an error someone has to explain.

Each of these can hold as many models as you like.

  • Models in your environment

    running inside your own environment

    Capacity you control completely, and the only one nothing leaves.

    Capacity you control

  • The provider you standardized on

    the agreement you already negotiated

    Capacity you are paying for whether or not you use it this month, so it gets used first.

    Used first, never replaced

  • Frontier providers

    over their APIs, as many as you keep keys for

    Your most capable models, and somewhere to go when the rest is full.

    Capacity to burst into

Tokenomics

Get the full picture.

Every call is recorded against the workflow that made it and the model that served it. Token use and spend across every provider on one screen, while there is still time to act on it.

Endpoint
nightshift-sr/auto
Overview|All workflows|Last 6 monthsRouting records
Calls routed, this month
5.24M
across four models
Rerouted around a limit
41.8K
no error reached a caller
P95 response
1.2s
held as volume tripled
Spend attributed
$140.7K
to the workflow that caused it
Volume against response timecallsp95 response
Where the calls ran4 destinations
Enterprise platform38%
Your environment31%
Frontier, primary21%
Frontier, overflow10%
Spend by workflow5 of 31, this month
WorkflowCallsCostOn frontierFinding
ops.schedule_replan742K$41,20061%Overpaying
finance.scenario_model18.4K$38,900100%Correct
quote.price_build196K$19,40044%Review
docs.invoice_extract2.14M$6,7002%Correct
support.order_status1.62M$4,1003%Correct
31 workflows5.24M$140,700

Sample data. Open Decisions, Token cost, Models, Workflows or Accuracy on the left for the same reporting from that seat.

  • Tokens tracked across every model

    One place for traffic and spend you currently read across separate invoices.

  • Attributed to a workflow

    Your invoices arrive by account. This arrives by the workflow that spent it.

  • Ahead of the bill

    You see a workflow getting more expensive with a month left to change it.

Governance

Some prompts should never leave.

The same reading that picks a model also rules models out. A prompt carrying personal data, or covered by a rule about where data may go, never reaches a provider that would break it.

  • Personal data never leaves your network.

    A prompt carrying it is held on a model inside your own perimeter.

  • A jailbreak never reaches a provider.

    Refused at the router, and the attempt is on the record.

  • Every decision can be reopened.

    What was read, what it matched, and which model answered it.

What changes

  • 96%

    lower effective cost

    Where routing let a small model recover most of a frontier model’s performance on a task.

    Published by the vLLM Semantic Router project

  • 86%

    answered without a metered call

    Share of prompts served by a free self-hosted model in the same published test.

    Published by the vLLM Semantic Router project

  • 3 of 3

    benchmarks matched or beaten

    A routed mixture against the single strongest model available, so the saving is not a quality trade.

    Published by the vLLM Semantic Router project

What we deliver

We deploy it and tune it around your workflows.

You get the deployment, the policies that decide where each request goes, and the evidence that those policies are right.

  • Deployed in your cloud

    It runs in your own cloud account, alongside the models it routes to.

  • We connect what you already run

    Starting with the provider you are committed to, then whatever you want beside it.

  • Policies built with you

    We start by watching, then write the rules once the records show what each workflow needs.

  • Standing evals

    Every routed workflow keeps being scored, so a routing decision stays defensible after we leave.

FAQs

  • No. Capacity you already pay for is the cheapest capacity you have, so the router sends traffic there first and anything else sits alongside it. No application changes, and neither does your agreement.

  • About 40 milliseconds at the median, measured by the router project on CPU with no dedicated GPU. That is 0.4 to 5% of the time the model itself takes.

  • It reads the prompt, rules out any model your policy does not allow for it, then scores what is left on quality, price and live speed. A slow or throttled model is not eligible until it recovers.

  • Nothing moves to a different model until the evals show it scores the same on that workflow’s own traffic. Where it does not, the workflow keeps the model it has.

  • Only where you allow it to. The router runs in your own cloud account, and a prompt carrying personal data or covered by a residency rule is held on a model inside your network.

  • Models you host yourself, whatever your cloud agreements include, and as many external providers as you keep keys for.

See where your model spend actually goes.

Start by watching. It records every request and routes none of them, so you begin with a picture of what you already run.