Throttle sits between your app and Claude/OpenAI APIs. It automatically routes requests to cheaper models when quality doesn't matter, caches repeated queries, and batches requests intelligently. no code changes needed. drop it in as a proxy. result: 70% cost cuts in production.
I built because we watched our inference bill hit 8k/month, optimized it, and realized every AI company has the same problem. open source on github."
Smart model routing based on task complexity / Intelligent query caching / Async request batching / Drop-in proxy deployment / Works with Claude and OpenAI / Open source, these three sections do 80% of the work. Description captures the founder story + problem + solution. Use Cases targets your buyer. Features show technical depth
AI startups bleeding money on inference costs. Teams using Claude/OpenAI at scale. Any company paying $1000+/month on LLM APIs and not optimizing.

shipping Throttle today. spent the last few months building this because teams are getting destroyed by LLM inference costs. the problem is real: you're paying for routes that don't need your best model, you're recomputing the same queries, your batching sucks. we cut our own spend from 8k to 2.4k just by being smart about it. open sourced it because if you're building with Claude or OpenAI at scale, you probably need this. grab it on github, test it on your actual traffic. would love to hear what breaks

shipping Throttle today. spent the last few months building this because teams are getting destroyed by LLM inference costs. the problem is real: you're paying for routes that don't need your best model, you're recomputing the same queries, your batching sucks. we cut our own spend from 8k to 2.4k just by being smart about it. open sourced it because if you're building with Claude or OpenAI at scale, you probably need this. grab it on github, test it on your actual traffic. would love to hear what breaks
Find your next favorite product or submit your own. Made by @FalakDigital.
Copyright ©2026. All Rights Reserved