Magic: AI Startup Claims Frontier-Level Pretraining for a Few Million Dollars
Set Trending Topics as a preferred source on Google.
After nearly two years of near silence, Magic is back with a research update. The AI startup founded by the two Viennese engineers Eric Steinberger and Sebastian De Ro has published a blog post that goes straight at one of the most expensive questions in the industry: how much compute does it actually take to train a competitive language model from scratch?
Magic’s answer is confident. Its pretraining recipe is now more than ten times more compute-efficient than the recipes behind the leading open-weight base models, the company says. The reasoning is blunt, and the team puts it in its own words:
“Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. Our new recipe matches DeepSeek V4 Pro’s pretrain using 50x less compute – that’s roughly half the FLOPs used for GPT3, or ~$0.5M on GB200.”
What Magic Is Claiming
Those 500,000 dollars refer to compute on Nvidia’s GB200 systems. In a second step, Magic says it scaled the same recipe up tenfold, to roughly four million dollars, and beat every publicly available base model on perplexity evaluations. Under DeepSeek V4 Pro’s recipe, the company argues, training a model of comparable capability would cost more than 100 million dollars, according to its own scaling laws.
The measurement does not come from the usual benchmark suites. Magic reports bits-per-byte loss on held-out data instead. The reason is technical: base models have not yet gone through reinforcement learning or fine-tuning, which makes them highly sensitive to prompt wording and makes sampling-based benchmarks unreliable. The held-out sets consisted of Magic’s own codebase, private code repositories acquired from other startups, recent low-citation research papers, and private math problems.
The Comparison Set
The choice of rivals is telling. Magic measures itself against open weights from China and the United States: DeepSeek V4 Pro and V4 Flash, Moonshot’s Kimi K2, and Nvidia’s Nemotron 3 Ultra. Base models from Anthropic, Google and OpenAI are missing for a simple reason, namely that they are not publicly available. As a proxy, Magic points to Kimi K3 and Meta’s Muse Spark, which it says indicate gains of 2.5x and 3.3x over Kimi K2.
To back up the numbers, the team says it recomputed the baseline figures across several inference engines and brought in the inference provider Fireworks for an independent check. One side finding: Nvidia’s Nemotron 3 outperforms DeepSeek V4 Pro across the board at the pretraining stage, which suggests the model’s weaker benchmark results may come down to its post-training rather than its base.
In a short reinforcement learning run on math problems, Magic’s current model reaches a pass rate of 72 percent, according to the charts in the blog post. For context, the company lists GPT-6 Astra at 100 percent, Claude Fable 5.1 at 98 percent and Kimi K3 at 75 percent, while noting itself that those models have presumably seen orders of magnitude more RL compute.
One caveat runs through all of it: every figure comes from the company, none has been independently verified, and Magic has still not shipped a model. The blog post closes with the line that the team looks forward to “releasing the thing”.
The People Behind the Startup
Magic was founded by Eric Steinberger and Sebastian De Ro, who met in a gifted-students program at the HTL Spengergasse in Vienna. As Trending Topics has reported in detail, the two spent their summer holidays training early AI models on school computers they had collected themselves. Steinberger was doing research at TU Wien at 19 and later worked at Trinity College Cambridge, MIT and Facebook AI. De Ro worked as a developer at Automic Software and the Austrian federal railways ÖBB, and later as CTO at the Austrian startup Firestart. The company is now headquartered in the United States.
Funding has been generous. A 23 million dollar round with Alphabet’s participation was followed by 117 million dollars and eventually a 320 million dollar round backed by former Google CEO Eric Schmidt, Jane Street, Sequoia, Atlassian and CapitalG, alongside investors Nat Friedman, Daniel Gross and Elad Gil. That valued the company at roughly 1.5 billion dollars. Magic also announced a partnership with Google Cloud to build compute capacity with tens of thousands of GB200 chips.
From Context Windows to Pretraining
What made Magic famous was a different claim. With LTM-2-mini, the startup presented a model with a context window of 100 million tokens, equivalent to about ten million lines of code or roughly 750 novels. At the time the company described that as 50 times what Google’s Gemini could handle and 780 times what OpenAI’s GPT-4o could take in. No publicly usable product has come out of it so far.
That sequence is exactly how Magic now describes its strategy. Pretraining, agentic RL and long context are together sufficient to build superhuman coding agents and automate AI research, the company writes. It started with long context, and pretraining comes next.
What Comes Next
The team names long-horizon reinforcement learning as its next focus, with agents that keep learning after deployment, along with work on alignment techniques. An updated “AGI Readiness Policy” covering deployment gates and safety requirements during RL training is in preparation.
The blog post doubles as a recruiting pitch. Magic is likely the smallest team in the world training trillion-parameter models, it says.

