Lives in your pull requests
Line annotations, a summary on the PR, and SARIF for GitHub code scanning.
Applied Benchmark runs cheaper and faster models against your code, prompts and evals. When one holds up, you get a pull request with the numbers. When it doesn't, you keep what you have.
example-model-0314 → example-model-0925
Retires in 23 days$1,764saved per month
Benchmarks measure general skill. Your prompts, your data and your edge cases decide which model is good enough for you.
Write evals, run both models, compare the output. It keeps slipping, and the expensive model stays.
Providers retire models every few months. Without numbers ready, you pick a replacement under a deadline and hope it holds.
Open source · Available now
Finds every model ID in your code and warns in CI before a provider retires one.
Early access
Tests cheaper and faster models on your evals, and opens a pull request only when quality holds.
For model providers
Launching a model? See how it performs inside real production codebases.
Line annotations, a summary on the PR, and SARIF for GitHub code scanning.
Sunset runs in your CI or your terminal. Your code stays there.
Each entry links to the provider's own notice, so you can check it.
No config, no sign-up. Or run it locally first.
- uses: AppliedBenchmark/sunset@v1$ npx @appliedbenchmark/sunset scan .We help teams that run language models in production find the cheapest, fastest model that still does the job, tested on their own code and evals. Sunset, our open-source scanner, is out today. Sunrise, which runs those tests, is what we're building next.
Yes. It's Apache-2.0, with no account, no API key and no usage limits. Sunrise will be a separate paid product, and Sunset will never depend on it.
No. Sunset reads your files locally, has no telemetry and doesn't use an LLM. Its only network call fetches the latest registry from GitHub. If you'd rather it made none, set registry: bundled.
OpenAI, Anthropic, Gemini API, Vertex AI, Amazon Bedrock, Azure OpenAI and Mistral, plus DeepSeek on a best-effort basis. OpenRouter slugs like openai/gpt-4.1 get the dates of the model behind them. Missing one you use? Adding it is a pull request.
Every date in the registry is quoted word for word from the provider's own page, with a link to it. CI checks each quote against a stored copy of that page, and a daily job opens a pull request when a provider changes theirs. A person reviews every one.
Only when it should. By default a run fails on models that are already retired or retire within 30 days, and warns from 90 days out. Model IDs in docs never fail a run. You can change all of this in .sunset.yml, or snooze a model with a reason and an end date.
A model chosen at run time, from an environment variable or a variable, can't be read from your code. Sunset lists it as unresolved and never fails a run on it. It checks the IDs in its registry, and IDs in documentation are listed as mentions only.
We run your current model and a candidate side by side on your own evals, inside your CI. If quality holds, you get a pull request with the change and the numbers behind it. If it doesn't, you get the numbers and no pull request.
We're setting early-access pricing together with our first teams. Join the list and tell us what it would be worth to you.
Almost there. Check your inbox and confirm your email.
You're on the list. We'll be in touch personally.
Something went wrong. Please try again, or email us at hello@appliedbenchmark.com.
By joining you agree that we store your email to contact you about early access. You can withdraw any time. Privacy policy.