Models#Benchmark#Model inference

Decision models are a category now: Jev

OpenAI and Amazon shipped Jev-style decision models within two weeks; TypeSafe reportedly talks a raise past $10B, three weeks after its $40M seed.

Latent Space podcast thumbnail: red-haired Jev founder Diogo Almeida beside the Jev product interface

In a Latent Space interview written up by QbitAI on October 3, TypeSafe AI founder Diogo Almeida — a former OpenAI researcher who worked on ChatGPT-adjacent efforts — laid out his company’s first product, Jev, as a “System-One model”: machine-native, large, programmable, unable to chat and built only to decide. Three API primitives map to the branch points of ordinary software — choice for switch statements, null for boolean tests (named after the Bernoulli distribution), score for ranking and thresholds — and the output goes straight back into the calling program, with no conversational surface at all. Training uses TypeSafe’s own RLCD method (reinforcement learning with closed-loop program verification as the target) rather than RLHF; the company deliberately ships no public benchmark, on the argument that leaderboards are trivially gameable. TypeSafe’s self-reported numbers: 38.7 million video views in six days after launch, and over a trillion tokens processed per day. The name comes from the Jevons paradox, and the pitch is extreme cost efficiency.

The founder: an emotional exit

Almeida’s background explains the product’s temperament. At OpenAI he worked on instructor-GPT-adjacent efforts, including replacing PPO data-cleaning pipelines with a homegrown algorithm; he describes leaving “with feelings” — seeing a direction and being unable to move it. He has recounted pitching his programmable-AI idea to Sam Altman and receiving encouragement in blunt terms: “this is so good, you should go do it.” TypeSafe describes itself as a data lab rather than a model lab: heavy on data hiring, deliberately refusing to train on user data (to avoid power-law overfitting), and aimed at dark-data analysis, coding agents, real-time decisions and games. The pricing claim circulating in coverage is $0.042 per million input tokens; the five-year goal is to move total factor productivity by 3%.

Two weeks from product to category

TechCrunch reported on September 30 that OpenAI shipped its own “Jev clone,” and again on October 1 that Amazon followed — under a headline declaring that “decision models flood the web.” A startup’s new paradigm being copied by two hyperscalers within a fortnight is rare in foundation models; the closest recent analogue is the rapid canonization of reasoning models after o1. TechCrunch’s reading of the OpenAI clone is the sharpest part: decision models may help frontier labs police their own swarming agents, because the expensive part of agentic planning is gratuitous generation, and branch decisions do not need generation at all.

The valuation rumor: 50x in three weeks

Multiple outlets report that TypeSafe is negotiating a new round above $1 billion at a valuation exceeding $10 billion — three weeks after emerging from stealth on September 15 with a $40 million seed led by DCVC (partner James Hardiman) at roughly $200 million. No terms or participants are confirmed, and a 50x jump in three weeks should be discounted accordingly. But the rumor and the fast-follow clones corroborate each other on one point: the market believes “models for programs” is a real category, not a pitch deck.

What to discount

Every headline number attached to Jev is company-reported: the trillion tokens a day, the claimed 444.6x cost advantage over an LLM-based workflow (which analysts have already questioned), and a “third-party JevBench v1.2.1” ranking at 75.3 that the founder himself disavows the relevance of. Almeida’s refusal to publish benchmarks makes his own numbers impossible to verify independently — a tension he at least addresses head-on: “Until you put the model into your own workflow and measure it against that workflow, you cannot know how it performs where it actually matters.” That shifts evaluation cost onto buyers, and doubles as an industry experiment in living without leaderboards.

The argument underneath

Almeida’s sharpest claim is economic: “almost all models today create commercial value close to zero.” Whether or not that is right, it lands in a week when an anonymous model topping usage charts was the biggest model story going — faith in leaderboards is being squeezed from both ends, by gaming and by the sense that they measure nothing you actually run. Decision models will not be settled by benchmarks either: the test is how many production switch statements and boolean branches really get replaced by choice, null and score. OpenAI and Amazon have placed their bets; check the workflows in three months.