Transcription
Almost every tech bro is currently losing their minds to this. Subq. It's a Miami startup that came out of nowhere claiming they built a model with a 12 million token context window. It's apparently 52 times faster than Flash Attention at 1/5th of the cost of current models.
Now, that sounds like a scam right off the bat, but let's look at the details. So, every model right now runs on Transformer Attention. This is your GPT, Claw, Gemini, all the other ones. The way it works is every single token in your context has to compare itself with every other token. And that sounds fine until you scale it. Meaning, if you double your context window, your cost quadruples. It's basically a quadratic increase between tokens and cost. It's why long context is really expensive and why no one has really pushed the model past a million tokens without some serious trade-offs.
What Subq claims they did is skip most of that work. They call it SSA, sub quadratic sparse attention. So let's say you drop an entire codebase and ask a question like why is this specific function breaking. The entire codebase is sitting in the context window, all 12 million tokens. The model sees everything. It just doesn't waste compute comparing every token against every other token. It reads the meaning of what you're asking first, then figures out what's actually related and only attends to that. So it goes straight to what actually matters. And what this means is if your context doubles, your cost also doubles. It doesn't quadruple.
Now, a lot of people suspect that SSA might just be a sparse attention model fine-tuned on top of an existing model like DeepSeek or Kimmy, which if true, this isn't really a ground-up architecture breakthrough. It's just an optimization. And keep in mind, there's no technical report. No one outside the company has independently tested any of this. And all of those benchmarks you might be hearing about are on the 1 million token model. The 12 million token headlines has zero results behind it.
Now, what is promising is that a lot of the team members come from Meta and Google, and the concept at least makes sense. They also have a waitlist, and of course, I'm on it. I'll let you know if anything comes out of this. Follow and I'll keep you posted.