Launch
Throttle
Visit
Example Image

Throttle

70% cheaper LLM inference. Drop in as a proxy

Visit

Throttle sits between your app and Claude/OpenAI APIs. It automatically routes requests to cheaper models when quality doesn't matter, caches repeated queries, and batches requests intelligently. no code changes needed. drop it in as a proxy. result: 70% cost cuts in production.

I built because we watched our inference bill hit 8k/month, optimized it, and realized every AI company has the same problem. open source on github."

Example Image
Example Image

Features

Smart model routing based on task complexity / Intelligent query caching / Async request batching / Drop-in proxy deployment / Works with Claude and OpenAI / Open source, these three sections do 80% of the work. Description captures the founder story + problem + solution. Use Cases targets your buyer. Features show technical depth

Use Cases

AI startups bleeding money on inference costs. Teams using Claude/OpenAI at scale. Any company paying $1000+/month on LLM APIs and not optimizing.

Comments

shipping Throttle today. spent the last few months building this because teams are getting destroyed by LLM inference costs. the problem is real: you're paying for routes that don't need your best model, you're recomputing the same queries, your batching sucks. we cut our own spend from 8k to 2.4k just by being smart about it. open sourced it because if you're building with Claude or OpenAI at scale, you probably need this. grab it on github, test it on your actual traffic. would love to hear what breaks

Premium Products
Social Links

Comments

shipping Throttle today. spent the last few months building this because teams are getting destroyed by LLM inference costs. the problem is real: you're paying for routes that don't need your best model, you're recomputing the same queries, your batching sucks. we cut our own spend from 8k to 2.4k just by being smart about it. open sourced it because if you're building with Claude or OpenAI at scale, you probably need this. grab it on github, test it on your actual traffic. would love to hear what breaks

Premium Products