← All stories
● Covered by 5 sources · 23 reportsMedium impact1 negative17 neutral4 positive

Laya, an open-source non-autoregressive AI model, released as alternative to Jev

🔄 Updated 18h ago — new reporting from The New Stack, Hacker News Front Page, TechCrunch
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Laya is an open-source, non-autoregressive AI model.
  • It offers schema-based decision-making using bidirectional encoders.
  • Laya runs in 32.8 ms on a single GPU, faster than Jev.
  • It supports over 100 languages and has Apache 2.0 weights.
  • Jev is TypeSafe AI's first "System One" model.
  • Jev is designed for programmatic statement evaluation and decision-making.
  • Diogo Almeida, ex-OpenAI engineer, co-founded TypeSafe AI and created Jev.
  • Jev is 194x faster and 445x cheaper than frontier AI models like GPT-6 Astra.
  • Jev uses Reinforcement Learning for Calibrated Decisions (RLCD) for training.
  • TypeSafe emerged from stealth after two years.
  • TypeSafe secured $40 million in seed funding led by DCVC.
  • Jev is a text-only AI model.
  • Jev provides structured, probabilistic decisions.
  • TypeSafe built a new architecture and sampler for Jev.
  • Diogo Almeida co-invented RLHF and InstructGPT at OpenAI.
  • A curation site for Jev launched, offering reviewed projects and code examples.
  • The Jev curation site includes 33 reviewed picks.
  • Jev was adopted faster than any other model in AI Gateway history, according to Vercel.
  • JevBench is a new benchmark for typed decision models.
  • JevBench evaluates models based on accuracy, latency, and cost.
  • JevBench allows configurable weighting of accuracy, latency, and price.
  • A full JevBench run asks 534 English decisions.
  • JevBench v1.3 score combines chance-corrected Intelligence, Calibration, Speed, and Cost.
  • Jev scored 74.4 on JevBench.
  • SemIf scored 73.1 on JevBench.
  • djeV scored 73.0 on JevBench.
  • Winnow-12B Q8 scored 71.2 on JevBench.
  • reflex 4B scored 70.3 on JevBench.
  • JevBench has a MIT harness, public items, frozen artifacts, scoring code, and public per-task outcomes.
  • JevBench has two no-signup demos.
  • JevBench is English-only.
  • JevBench latency is measured from one German server.
  • JevBench applies a disclosed x2 adjustment (+150 ms) for local/demo latency.
  • Jev returns decisions, scores, or choices with probabilities.
  • Jev is designed for machine consumption.
  • Jev has manipulation vulnerabilities similar to large language models.
  • Attackers can alter Jev's verdicts with minimal effort and cost.
  • Changing Jev's mind cost about 50 cents in testing.
  • Jev answers structured questions about an input in a single forward pass.
  • Jev provides epistemically honest probabilities.
  • Jev's speed comes from its non-autoregressive output.
  • Jev emits the entire structured answer at once.
  • Jev has an end-to-end latency of 70–500 ms.
  • Jev is 40x–200x faster than frontier LLMs on equivalent tasks.
  • Jev uses typed questions instead of prompts.
  • A Python wrapper allows querying LLMs and vision models.
  • The wrapper forces single-token responses and analyzes log probabilities.
  • The method offers flexibility for condition changes via plain text.
  • The wrapper extends beyond text-only inputs to include image attachments for vision models.
  • OpenJev and SemIf are self-hostable projects appearing around Jev.
  • The method uses 'max_completion_tokens': 1 and 'logprobs': true in Chat Completions requests.
  • The LLM API returns the letter plus the model's log probabilities for alternative tokens.
  • Processing the input still costs time.
  • A shared state prefix can be KV-cached if the backend supports it.
  • Jev's documented request format currently describes only text/JSON state.
  • GLM-5.3-Flash can be configured to make typed decisions with probabilities in a single forward pass.
  • The method allows LLMs to be used for high-volume decision-making tasks.
  • Johannes Hötter is VP Growth.
  • Marko Rosenmüller, PhD, is Technical Lead AI.
  • The approach was evaluated using GLM-5.3-Flash running on Privatemode.
  • The setup with GLM-5.3-Flash enables typed decisions on images.
  • Jev is criticized for unreliable outputs.
  • Jev requires users to build their own evaluation and ground-truth pipelines.
  • Jeff is a new 0.8B parameter decision model.
  • Jeff is fine-tuned from Qwen3.5 and Gemma 4.
  • Jeff runs in 22 ms on an RTX PRO 6000.
  • Jeff runs in 28 ms on an Apple M4 Max (MLX).
  • Jeff supports zero-shot classification tasks.
  • Jeff can be fine-tuned on custom examples to improve accuracy.
  • Jeeves is a new AI model.
  • Jeeves is based on Qwen3.5-9B.
  • Jeeves integrates a reasoning step before making decisions.
  • Jeeves outperforms Jev and Kev models in accuracy.
  • Jeeves was trained with SFT and CISPO.
  • Jeeves includes a block-4 diffusion drafter.
  • Jeeves beats Kev-9B and Jev on test data (0.889 vs 0.822 and 0.857).
  • Jeeves beats Jev on JevBench's public tiers (0.935 vs 0.866).
  • Jeeves supports yes/no, multiple-choice, and rating questions.
  • Jeeves has a Jev-compatible API.
  • Jeeves takes 0.3 seconds per request without thinking on an H100.
  • Jeeves takes 3.3 seconds median with thinking on an H100.
  • Jeeves runs on CUDA (Hopper for FP8 kernel).
  • Jevstiller is a local AI model that replicates Jev's text classification answers.
  • Jevstiller achieves 98% agreement with Jev.
  • Jevstiller reduces response time from 300ms to 15ms.
  • Jevstiller offloads requests from the Jev API.
  • Jevstiller was developed in September 2026.
  • Jevstiller uses a frozen sentence encoder (bge-small, 384 dimensions, ONNX Runtime on CPU).
  • Jevstiller trains a multinomial logistic regression head.
  • OpenAI announced its new Decisions API at its annual DevDay conference.
  • OpenAI's Decisions API is built on its Luna model.
  • OpenAI's Luna model is the smallest and most affordable in its lineup.
  • OpenAI's Decisions API is available in limited preview.
  • OpenAI plans a broad release of the Decisions API in the coming days.
  • Jev is being explored for integration with Apache Arrow.
  • Jev's outputs currently use JSON.
  • Featherless launched Simple Jev.
  • Simple Jev is an open-source library.
  • Simple Jev converts open-source AI models into high-speed, zero-shot classification engines.
  • Simple Jev extends structured decision functionality to open-source models.
  • Simple Jev has image capabilities on Featherless's hosted endpoints.
  • Eugene Cheah is CEO and co-founder of Featherless.
  • OpenAI's Decisions API provides functionality similar to TypeSafe AI's Jev model.
  • OpenAI's Decisions API allows the Luna model to choose from predefined options.
  • OpenAI's Decisions API aims for faster and cheaper operations compared to general LLMs.
  • Sam Altman announced OpenAI's Decisions API.
  • OpenAI's Decisions API allows the Luna model to classify images.
  • OpenAI's Decisions API allows the Luna model to choose different agent behaviors.
  • Diogo Almeida joked on X about OpenAI's Decisions API.
  • Jev processes structured data and questions in parallel.
  • Jev returns Choice, Score, and Noul answers.
  • Jev provides confidence values with its answers.
  • Jev input costs $0.042 per million tokens.
  • Jev output is free.
  • Jev has a context window of 32,000 tokens.
  • Jev reached nearly 13% of paid teams on Vercel within 24 hours.
  • Jev's adoption on Vercel was twice the share of the GPT-5.6 family.
  • Netlify added Jev.
  • LangChain shipped a TypeSafeClassifier integration with model routing.
  • LangChain shipped an AutoMode middleware that screens tool calls.
  • Five independent Elixir clients for Jev appeared within days.
  • AWS launched Strands Decider 2B, a local decision model.
  • Strands Decider 2B uses Qwen3.5-2B as its base model.
  • Strands Decider 2B replaces the language-model head with a pointer head.
  • Cloudflare released Clef and Clef-flash, open-source decision models.
  • Clef and Clef-flash are hosted on Workers AI.
  • Clef is the leader on the Jev Decision Index.
  • Clef and Clef-flash are Jev-API compatible.
  • Clef and Clef-flash are open-sourced on Hugging Face under Apache 2.0 license.

Laya's Release and Purpose

Laya, a new open-source AI model, has been released as an alternative to TypeSafe AI's Jev. Laya is designed for lightning-fast probability predictions over structured schemas, distinguishing itself from traditional autoregressive and generative text models. It aims to provide an open and efficient solution for System 1 reflex decisions in AI pipelines.

Technical Specifications and Performance

Laya is built on bidirectional encoders, allowing it to achieve inference times of 32.8 milliseconds on a single GPU, or 7.2 milliseconds per question when batched. This performance is stated to be 6 to 8 times faster than Jev. Laya also supports over 100 languages and is released with 100% open-source Apache 2.0 weights, eliminating API subscription costs.

Background and Comparison to Jev

The developer of Laya previously published research on non-autoregressive, reinforcement learning-guided schema-based decision systems in March and October 2025. TypeSafe AI, founded by Diogo Almeida, later launched Jev in September 2026, proposing a similar non-autoregressive decision concept. Jev uses RLCD (Reinforcement Learning for Calibrated Decisions) for confidence distributions and schema choices, charging $0.042 per million input tokens with typical response times around 150 ms. Unlike Laya, Jev launched without public technical papers, open weights, or open training datasets.

Architectural Approach

Laya addresses architectural limitations of earlier models by focusing on a completely open, horizontal System 1 decision model family. The core idea behind Laya, and similar models, is to move away from using generative LLMs for simple reflex decisions, which are identified as a bottleneck in modern AI pipelines.

Updates

🕒 2026-10-01 · new reporting from The New Stack, Hacker News Front Page, TechCrunch
  • AWS launched Strands Decider 2B, a local decision model.
  • Strands Decider 2B uses Qwen3.5-2B as its base model.
  • Strands Decider 2B replaces the language-model head with a pointer head.
  • Cloudflare released Clef and Clef-flash, open-source decision models.
  • Clef and Clef-flash are hosted on Workers AI.
  • Clef is the leader on the Jev Decision Index.
  • Clef and Clef-flash are Jev-API compatible.
  • Clef and Clef-flash are open-sourced on Hugging Face under Apache 2.0 license.
🕒 2026-10-01 · new reporting from InfoQ
  • Jev processes structured data and questions in parallel.
  • Jev returns Choice, Score, and Noul answers.
  • Jev provides confidence values with its answers.
  • Jev input costs $0.042 per million tokens.
  • Jev output is free.
  • Jev has a context window of 32,000 tokens.
  • Jev reached nearly 13% of paid teams on Vercel within 24 hours.
  • Jev's adoption on Vercel was twice the share of the GPT-5.6 family.
  • Netlify added Jev.
  • LangChain shipped a TypeSafeClassifier integration with model routing.
  • LangChain shipped an AutoMode middleware that screens tool calls.
  • Five independent Elixir clients for Jev appeared within days.
🕒 2026-09-30 · new reporting from TechCrunch
  • OpenAI's Decisions API provides functionality similar to TypeSafe AI's Jev model.
  • OpenAI's Decisions API allows the Luna model to choose from predefined options.
  • OpenAI's Decisions API aims for faster and cheaper operations compared to general LLMs.
  • Sam Altman announced OpenAI's Decisions API.
  • OpenAI's Decisions API allows the Luna model to classify images.
  • OpenAI's Decisions API allows the Luna model to choose different agent behaviors.
  • Diogo Almeida joked on X about OpenAI's Decisions API.
🕒 2026-09-29 · new reporting from Hacker News Front Page, The New Stack
  • Jev is being explored for integration with Apache Arrow.
  • Jev's outputs currently use JSON.
  • Featherless launched Simple Jev.
  • Simple Jev is an open-source library.
  • Simple Jev converts open-source AI models into high-speed, zero-shot classification engines.
  • Simple Jev extends structured decision functionality to open-source models.
  • Simple Jev has image capabilities on Featherless's hosted endpoints.
  • Eugene Cheah is CEO and co-founder of Featherless.
🕒 2026-09-29 · new reporting from The New Stack
  • OpenAI announced its new Decisions API at its annual DevDay conference.
  • OpenAI's Decisions API is built on its Luna model.
  • OpenAI's Luna model is the smallest and most affordable in its lineup.
  • OpenAI's Decisions API is available in limited preview.
  • OpenAI plans a broad release of the Decisions API in the coming days.
🕒 2026-09-29 · new reporting from Hacker News Front Page
  • Jevstiller is a local AI model that replicates Jev's text classification answers.
  • Jevstiller achieves 98% agreement with Jev.
  • Jevstiller reduces response time from 300ms to 15ms.
  • Jevstiller offloads requests from the Jev API.
  • Jevstiller was developed in September 2026.
  • Jevstiller uses a frozen sentence encoder (bge-small, 384 dimensions, ONNX Runtime on CPU).
  • Jevstiller trains a multinomial logistic regression head.
🕒 2026-09-29 · new reporting from Hacker News Front Page
  • Jeeves is a new AI model.
  • Jeeves is based on Qwen3.5-9B.
  • Jeeves integrates a reasoning step before making decisions.
  • Jeeves outperforms Jev and Kev models in accuracy.
  • Jeeves was trained with SFT and CISPO.
  • Jeeves includes a block-4 diffusion drafter.
  • Jeeves beats Kev-9B and Jev on test data (0.889 vs 0.822 and 0.857).
  • Jeeves beats Jev on JevBench's public tiers (0.935 vs 0.866).
  • Jeeves supports yes/no, multiple-choice, and rating questions.
  • Jeeves has a Jev-compatible API.
  • Jeeves takes 0.3 seconds per request without thinking on an H100.
  • Jeeves takes 3.3 seconds median with thinking on an H100.
  • Jeeves runs on CUDA (Hopper for FP8 kernel).
🕒 2026-09-28 · new reporting from Hacker News Front Page
  • Jeff is a new 0.8B parameter decision model.
  • Jeff is fine-tuned from Qwen3.5 and Gemma 4.
  • Jeff runs in 22 ms on an RTX PRO 6000.
  • Jeff runs in 28 ms on an Apple M4 Max (MLX).
  • Jeff supports zero-shot classification tasks.
  • Jeff can be fine-tuned on custom examples to improve accuracy.
🕒 2026-09-27 · new reporting from Hacker News Front Page
  • Jev is criticized for unreliable outputs.
  • Jev requires users to build their own evaluation and ground-truth pipelines.
🕒 2026-09-27 · new reporting from Hacker News Front Page
  • GLM-5.3-Flash can be configured to make typed decisions with probabilities in a single forward pass.
  • The method allows LLMs to be used for high-volume decision-making tasks.
  • Johannes Hötter is VP Growth.
  • Marko Rosenmüller, PhD, is Technical Lead AI.
  • The approach was evaluated using GLM-5.3-Flash running on Privatemode.
  • The setup with GLM-5.3-Flash enables typed decisions on images.
🕒 2026-09-26 · new reporting from Hacker News Front Page
  • A Python wrapper allows querying LLMs and vision models.
  • The wrapper forces single-token responses and analyzes log probabilities.
  • The method offers flexibility for condition changes via plain text.
  • The wrapper extends beyond text-only inputs to include image attachments for vision models.
  • OpenJev and SemIf are self-hostable projects appearing around Jev.
  • The method uses 'max_completion_tokens': 1 and 'logprobs': true in Chat Completions requests.
  • The LLM API returns the letter plus the model's log probabilities for alternative tokens.
  • Processing the input still costs time.
  • A shared state prefix can be KV-cached if the backend supports it.
  • Jev's documented request format currently describes only text/JSON state.
🕒 2026-09-25 · new reporting from Hacker News Front Page
  • Jev answers structured questions about an input in a single forward pass.
  • Jev provides epistemically honest probabilities.
  • Jev's speed comes from its non-autoregressive output.
  • Jev emits the entire structured answer at once.
  • Jev has an end-to-end latency of 70–500 ms.
  • Jev is 40x–200x faster than frontier LLMs on equivalent tasks.
  • Jev uses typed questions instead of prompts.
🕒 2026-09-24 · new reporting from Hacker News Front Page
  • Jev returns decisions, scores, or choices with probabilities.
  • Jev is designed for machine consumption.
  • Jev has manipulation vulnerabilities similar to large language models.
  • Attackers can alter Jev's verdicts with minimal effort and cost.
  • Changing Jev's mind cost about 50 cents in testing.
🕒 2026-09-23 · new reporting from Hacker News Front Page
  • JevBench is a new benchmark for typed decision models.
  • JevBench evaluates models based on accuracy, latency, and cost.
  • JevBench allows configurable weighting of accuracy, latency, and price.
  • A full JevBench run asks 534 English decisions.
  • JevBench v1.3 score combines chance-corrected Intelligence, Calibration, Speed, and Cost.
  • Jev scored 74.4 on JevBench.
  • SemIf scored 73.1 on JevBench.
  • djeV scored 73.0 on JevBench.
  • Winnow-12B Q8 scored 71.2 on JevBench.
  • reflex 4B scored 70.3 on JevBench.
  • JevBench has a MIT harness, public items, frozen artifacts, scoring code, and public per-task outcomes.
  • JevBench has two no-signup demos.
  • JevBench is English-only.
  • JevBench latency is measured from one German server.
  • JevBench applies a disclosed x2 adjustment (+150 ms) for local/demo latency.
🕒 2026-09-22 · new reporting from Hacker News Front Page
  • A curation site for Jev launched, offering reviewed projects and code examples.
  • The Jev curation site includes 33 reviewed picks.
  • Jev was adopted faster than any other model in AI Gateway history, according to Vercel.
🕒 2026-09-21 · new reporting from The New Stack
  • TypeSafe emerged from stealth after two years.
  • TypeSafe secured $40 million in seed funding led by DCVC.
  • Jev is a text-only AI model.
  • Jev provides structured, probabilistic decisions.
  • TypeSafe built a new architecture and sampler for Jev.
  • Diogo Almeida co-invented RLHF and InstructGPT at OpenAI.
🕒 2026-09-21 · new reporting from Tom's Hardware
  • Jev is TypeSafe AI's first "System One" model.
  • Jev is designed for programmatic statement evaluation and decision-making.
  • Diogo Almeida, ex-OpenAI engineer, co-founded TypeSafe AI and created Jev.
  • Jev is 194x faster and 445x cheaper than frontier AI models like GPT-6 Astra.
  • Jev uses Reinforcement Learning for Calibrated Decisions (RLCD) for training.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Amazon Web Services (AWS) released Strands Decider 2B, an open-source decision model designed for high-speed, low-cost sorting of pre-decided options with confidence scores. This model addresses the need for AI intelligence more suited to computer automation than large language models (LLMs) for specific agentic workflows.

Cloudflare has released Clef and Clef-flash, two open-source decision models hosted on Workers AI, along with a new reinforcement learning (RL) product for fine-tuning. These models provide bounded structured outputs for decision-making workflows and are compatible with the Jev-API.

AWS introduced Strands Decider 2B, a new local decision model designed to select from developer-supplied options or return numerical scores. This model provides faster decisions and confidence scores by restricting the answer space, making it useful for routing natural language requests and evaluating agent actions.

TypeSafe AI, founded by a former OpenAI researcher, launched Jev, an AI model that provides typed, probabilistic decisions instead of text generation. This model processes structured data and questions in parallel, returning Choice, Score, and Noul answers with probability distributions and confidence values. Jev's design allows software to act directly on its outputs, offering faster performance and lower costs compared to traditional LLMs for specific tasks.

OpenAI announced its new Decisions API at Dev Day, which provides functionality similar to TypeSafe AI's Jev model for software automation. This API allows OpenAI's Luna model to choose from predefined options, aiming for faster and cheaper operations compared to general LLMs.

Featherless launched Simple Jev, an open-source library that converts open-source AI models into high-speed, zero-shot classification engines. This tool allows applications to evaluate data and return categorical assignments or binary choices without generating conversational text, aiming to provide faster and more cost-effective AI decision-making.

OpenAI announced its new Decisions API, built on its Luna model, at its annual DevDay conference. This API provides predefined answers with confidence scores, offering a faster and more accurate alternative to traditional LLMs for classification and routing tasks.

TypeSafe AI's Jev model, designed to convert natural language into typed decisions with scores and probabilities, is being explored for integration with Apache Arrow. This integration aims to improve the efficiency of data pipelines by using Arrow's structured data capabilities instead of JSON for Jev's outputs. The goal is to enable more sophisticated probabilistic workflows by combining Jev's decision-making with Arrow's performance benefits.

Jevstiller developed a local AI model that replicates Jev's text classification answers with 98% agreement, reducing response time from 300ms to 15ms. This allows for faster decision-making in applications like agent loops or game ticks by offloading most requests from the Jev API.

The new Jeeves AI model, based on Qwen3.5-9B, integrates a reasoning step before making decisions, outperforming previous Jev and Kev models in accuracy. This development provides a more reliable AI for calibrated decision probabilities, reducing the need for human fallback in AI pipelines.

Jeff released new 0.8B and 2B parameter decision models, fine-tuned from Qwen3.5 and Gemma 4, for zero-shot classification tasks. These models offer fast, calibrated probability outputs for various options, operating locally with low latency.

The AI model Jev, developed by TypeSafe AI, is criticized for its unreliable outputs and the necessity for users to build their own evaluation and ground-truth pipelines. This raises concerns about the practical utility of AI products that require significant user effort to validate their performance.

Researchers demonstrated that the GLM-5.3-Flash LLM can be configured to make typed decisions with associated probabilities in a single forward pass, similar to specialized decision models like Jev. This method allows LLMs to be used for high-volume decision-making tasks, overcoming previous limitations in speed and cost.

A new Python wrapper allows querying Large Language Models (LLMs) and vision models by forcing single-token responses and analyzing log probabilities. This method offers flexibility for condition changes via plain text, extending beyond text-only inputs to include image attachments for vision models.

TypeSafe AI launched Jev, a new "System One model" designed for structured question answering with probabilistic outputs, rather than conversational AI. The model emphasizes calibration over raw accuracy, addressing a common limitation in production classifiers by providing epistemically honest probabilities.

TypeSafe AI released Jev, a new AI model that provides decisions, scores, or choices with probabilities, designed for machine consumption. Testing revealed that Jev, despite its different architecture, shares similar manipulation vulnerabilities to large language models, allowing attackers to alter its verdicts with minimal effort and cost.

A new dependency upgrade agent, safe-upgrade, has been developed using Jev, LangGraph, and Tenuo. This agent aims to automate and improve the safety of dependency upgrades by separating judgment, control flow, and authority in the process.

An analysis suggests OpenAI could replicate TypeSafe's Jev large language model, which has seen rapid adoption. OpenAI's existing LLM capabilities and potential to integrate Jev-like features into its models could allow it to offer similar functionality.

A new curation site, Jev, has launched, offering reviewed projects, reusable skills, and code examples for TypeSafe AI's Jev model. The site provides a starting point for developers to explore Jev's capabilities in answering typed questions about text or JSON for bounded decisions.

JevBench, a new benchmark, has been released to evaluate typed decision models based on accuracy, latency, and cost. It aims to provide a standardized way to compare the performance of models that return bounded choices and probabilities.

TypeSafe emerged from stealth with $40 million in seed funding to launch Jev, a text-only AI model designed for machines to make decisions within software applications. Jev offers structured, probabilistic decisions, contrasting with sequential LLMs that TypeSafe claims are inefficient for computer use.

TypeSafe AI, co-founded by ex-OpenAI engineer Diogo Almeida, launched Jev, a new "System One" AI model designed for programmatic statement evaluation and decision-making. Jev claims to be significantly faster and cheaper than traditional LLMs by focusing on structured outputs and parallel processing, offering an alternative for specific AI applications.

Laya, an open-source, non-autoregressive AI model, has been released as a faster alternative to TypeSafe AI's Jev. Laya offers schema-based decision-making with faster inference times and no API costs, contrasting with Jev's proprietary model.