Loading...
Loading...
Luna got 80 percent cheaper and Terra dropped 20 percent overnight. Plus the wild detail that Sol helped optimize its own serving kernels. I reworked my bot economics the same day.
Every few months there's a price drop that makes you rethink what's worth automating. This was one of those.
$0.20 per million input tokens2.5x faster at double priceTwenty cents. When my alerts bot started last year, that kind of capability cost fifty times more, and I budgeted around it like a utility bill. The subscription math changed too: Terra and Luna usage now burns fewer credits inside ChatGPT Work and Codex, so the same plan quietly got more generous.
Buried in their efficiency post from July 29: Sol autonomously rewrote production CUDA kernels, ran hundreds of experiments on token generation, and cut serving costs by 20 percent. A human-led process set the guardrails, but the model did the optimizing, monitored its own training runs, and intervened when things drifted.
A model improving the infrastructure that serves itself. People will argue about what to call that for years.
Whatever you call it, the practical effect showed up in my invoice within a day. That's the part worth sitting with. Research demos stay in the lab, but kernel optimization compounds straight into everyone's pricing.
GPT-5.6 guarantees 30 minute cache lifeMonthly spend dropped by more than half and nothing got slower in any way I can measure.
Quick follow-up since I've lived with this setup a bit: the Luna classification tier hasn't misfired once, cache hits are saving another chunk on repetitive prompts, and the only surprise was pleasant. My alert volume doubled during a volatile stretch and the bill barely moved.
The lesson I keep relearning: reprice your stack every time a lab sneezes. Assumptions that made sense in June were obsolete by Friday.
Cheers, Yassen