Laya: free System One decisions for your agents, built into Melaya and reachable over MCP
Laya is an open System One decision model built into Melaya. Free typed decisions with zero output tokens in agents, pipelines, Fast Mode and MCP.
Laya is an open-weight System One model, released under Apache-2.0, that answers typed questions about a piece of text and returns calibrated probabilities with zero generated tokens. Its authors publish it as the pip package laya and on Hugging Face. Melaya did not build Laya. Melaya runs it on its own CPU and builds it in for free: the decide tools, the canvas decide step, event triggers, and Fast Mode browser and phone agents, including from Claude or ChatGPT over MCP.
What Laya is, and who publishes it
Laya is a System One model: a fast decision model that reads a state, such as an email, a post, a page or a list row, and answers typed questions about it. It returns a calibrated probability for each option and generates zero tokens. It decides. It does not write.
Laya is open-weight under the Apache-2.0 license. NandhaKishorM publishes the code as the pip package laya, and convaiinnovations publishes the weights on Hugging Face. Melaya did not create the model. Melaya runs the open Laya model on its own servers and builds it into its agents and pipelines, so you get it with no setup and no bill.
Under the hood, Laya is a ModernBERT encoder with small decision heads. The English base checkpoint has 421M parameters, and a multilingual checkpoint also exists. Every question takes one of three shapes.
- choice: pick one option from a set, such as billing, technical or sales
- score: place the state on 2 to 10 ordered levels, such as low, medium or high
- noul: a yes or no answer where the probability itself is the signal
- Every answer carries answer_confidence, the calibrated probability of the winning option
System One vs an LLM: when to use which
A large language model is System Two. It reasons, writes and calls tools, and you pay for every token it generates. Laya is System One. It applies the same typed judgment to thousands of items, one encoder pass per question, with nothing to parse afterwards.
Use an LLM to plan, to write and to handle open-ended reasoning. Use Laya when the job is to apply one judgment to every item: route a ticket, screen a queue, check a list row, gate a trigger. Good pipelines use both. Laya filters, and the LLM spends its tokens on the few items that survive.
Know its limits. Laya is reliable on concrete surface questions, such as whether a reply is automatic or whether a message asks about pricing. It is weaker on nuanced relevance, and its authors describe the base checkpoints as a fast base to specialise. A 2-option choice beats a bare yes or no, and score is the weakest of the three types.
How Melaya runs Laya for free
Melaya runs Laya in laya-local, an internal service that keeps the model warm on Melaya's own CPU. Measured on the production server, a single decision takes about 105 ms on the bf16 runtime, and batches run at about 69 ms per decision. That is CPU speed, not GPU speed, and you pay nothing per decision.
Agents reach it through built-in decide tools. They are core tools: no connector, no API key, no Connect step. Per-user rate limits keep the shared CPU available to everyone.
- decide: ask one state a set of typed questions in one pass
- decide_check: a one-shot gate that returns a single crisp answer
- decide_batch: score many items against the same questions and keep the top results
- decide_batch_file: read the items from a JSON file first, for example a harvest of posts
- The canvas decide step: a pipeline step with no LLM call that can stop the run when an answer says so
- Event triggers: a decide gate can screen each incoming event before any action runs
Laya inside Fast Mode and over MCP
Fast Mode drives a browser or an Android phone with far fewer LLM round trips. The LLM reads the screen once and sends several steps in plain words. Melaya finds each target, acts through the normal path, checks the effect and hands the screen back at the first surprise.
Deterministic code grounds each target first. Laya is the gated fallback: when a step describes its target loosely, Laya picks among a short list of candidates, and Melaya acts only when the answer clears a confidence bar. Laya also checks conditions on list rows, such as whether a row is a partner at a given firm.
Over MCP, Claude, ChatGPT, Cursor and Claude Code reach Laya in two ways: the Fast Mode tools, and the pipelines you run that use the decide tools or a decide step. There is no standalone laya MCP tool. In practice, Laya MCP means Laya working inside the tools of the Melaya MCP Server.
- MCP: melaya_browser_fast and melaya_phone_fast
- Pipelines and the Assistant: browser_fast and phone_fast
- Your rules still apply: allowed sites, approval cards and the kill switch
- Phone control is Android only
Where Melaya itself uses Laya
Melaya runs its own features on the same free engine. Backlink submissions in Melaya Marketing fill directory forms with Fast Mode, so Laya grounds loosely described fields when exact matching finds nothing.
The whole-site SEO audit has a Laya judgment layer for per-page questions. A Laya answer raises an issue only once its question passes a validated precision threshold. Until then, the audit relies on deterministic checks and never raises an issue from an unvalidated Laya answer.
Melaya's own outreach pipelines use Laya as a reply guard. A 2-option choice spots replies that ask about money terms and routes them to a human.
Laya vs Jev (TypeSafe)
Jev is the hosted System One model from TypeSafe. It runs on GPUs, answers in milliseconds and is metered per token against your own API key. Laya runs on Melaya's CPU, takes about 105 ms per decision and costs nothing.
Both take the same schema: one state, typed questions, the same three types. A workflow moves from the jev_* tools to the decide_* tools without rewriting a question. They are different models with different calibration, so re-check your thresholds when you switch.
- Pick Laya for volume, for cost and for decisions that stay on Melaya's own hardware
- Pick Jev for the lowest latency, low volumes and no shared CPU
Get started with Laya
Create a free Melaya account. The decide tools are built in: the Assistant always has them, and you add them to any agent with nothing to connect.
- Add decide_batch to an agent, or drop a decide step on the canvas
- Write a labelled state with the judged text first. The English base reads about 1,100 characters of state
- Ask a 2-option choice with short keyword descriptions instead of a bare yes or no
- Gate on answer_confidence and fit the threshold on a few labelled items
- From Claude or ChatGPT, connect the Melaya MCP Server, then run your pipeline or let the Fast Mode tools call Laya for you
Frequently asked questions
Is Laya made by Melaya?
No. Laya is an open-weight Apache-2.0 model from its upstream authors, with the code on pip as laya and the weights on Hugging Face from convaiinnovations. Melaya runs the open model on its own CPU and builds it into agents, pipelines and Fast Mode for free.
Can I use Laya from Claude or ChatGPT via MCP?
Yes, through the Melaya MCP Server. The melaya_browser_fast and melaya_phone_fast tools use Laya as a gated fallback, and any pipeline you run that uses the decide tools or a decide step runs Laya. There is no separate laya tool to call.
Is Laya free in Melaya?
Yes. The decide tools, the canvas decide step and Fast Mode run Laya on Melaya's CPU with no per-decision charge and no API key. Per-user rate limits keep the shared engine fair.
How fast is Laya?
On Melaya's CPU, a single decision takes about 105 ms, and batched decisions about 69 ms each. Long states cost more, so keep the judged text short. GPU-backed engines such as Jev answer in milliseconds.
What is the difference between Laya and Jev?
Both are System One models with the same question schema. Jev is TypeSafe's hosted engine, metered per token on GPUs. Laya is open-weight and runs free on Melaya's CPU.
Can Laya write text or replace my LLM?
No. Laya only decides: it picks an option, scores a level or answers yes or no, with zero generated tokens. Keep the LLM for planning, writing and nuanced judgment.
