
Cloudflare Open Sources Decision Models for AI Agents
Quick Answer
Cloudflare has open-sourced Clef, a set of AI decision models with 9B and 27B parameters, designed for specific decision-making tasks.
Quick Take
The 27B model processes multimodal inputs and boasts a 64k context window, while the 9B model targets latency-sensitive applications with a median latency of 38.8 ms, significantly faster than the larger model's 209.3 ms.
Key Points
- Clef models can process text, JSON, images, and video in a single forward pass.
- The 27B Clef model has a 64k context window, doubling Jev's capacity.
- Clef-Flash, the 9B model, achieves median latency of 38.8 ms, much faster than the 27B model.
- Cloudflare plans to fine-tune Clef for specific use cases using historical data.
- Models are available on Hugging Face for developers to experiment and provide feedback.
DeepSignal Analysis
What happened
Cloudflare has open-sourced Clef, a set of AI decision models with 9B and 27B parameters. The 27B model can process multimodal inputs and has a 64k context window, while the 9B model is optimized for low latency, achieving a median of 38.8 ms compared to 209.3 ms for the larger model.
Key evidence
- Cloudflare released Clef, which includes 9B and 27B parameter models designed for decision-making tasks rather than text generation.
- The Clef-Flash model, with 9B parameters, has a median latency of 38.8 ms, significantly faster than the 27B model's 209.3 ms.
- Cloudflare plans to fine-tune Clef for specific applications like support triage, using historical data to enhance its performance.
Why it matters
The introduction of Clef models represents a shift towards specialized AI decision-making tools that can handle multimodal inputs. This could enhance operational efficiency in various applications, particularly in customer support and automation. However, the effectiveness of these models in real-world scenarios remains to be fully evaluated, especially concerning their accuracy and reliability in critical decision-making contexts.
What to watch
📖 Reader Mode
~3 min readDuring its recent "Birthday Week", Cloudflare announced Clef, a set of open-weight AI models designed to choose between predefined options rather than generate text. Cloudflare released 9B- and 27B-parameter models, along with a platform for adapting them to specific decision-making tasks.
Clef is a 27B multimodal model that takes a state and a schema of typed questions as input and returns decisions. It can process text, JSON, images, or video, and returns a probability for each allowed option for every question in a single forward pass, without generating free-form text or requiring output parsing.

Source: Cloudflare blog
A decision model classifies inputs and returns typed outcomes with probabilities, allowing AI agents to use these results to determine how to act, such as routing a support request, escalating it, or deferring the decision to a human. Cloudflare’s API is compatible with the popular Typesafe AI’s Jev System One model.

Source: Cloudflare blog
Michelle Chen, group product manager at Cloudflare, Alex Reneau, principal machine learning engineer at Cloudflare, and Kevin Flansburg, senior engineering manager at Cloudflare, write:
Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action.
The hyperscaler also released Clef-Flash, a smaller 9B multimodal model designed for latency-sensitive decisions. It has a median latency of 38.8 ms in Cloudflare’s benchmarks, compared with 209.3 ms for the 27B Clef model. Chen, Reneau, and Flansburg explain how Clef differs from other decision models:
First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against.
On a popular Hacker News thread, the community discusses decision models, latency, open weights, and whether this is really a new model category. Jacek Złydach writes:
It's not a ‘new paradigm’, it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it.
The benchmark results raised further questions, with user SebastianSosa warning:
Public benchmarks are easy to cheat, if I am Typesafe, I would also release a public benchmark to distract otherwise competent people from overfitting to a benchmark instead of making something actually useful.
Cloudflare plans to fine-tune Clef for specific use cases, such as support triage and bot classification, using its historical labelled data to improve accuracy and speed. On Reddit, user bugra_sa writes:
I'd care more about whether Clef knows when to punt than its raw accuracy score. Test it on cases where a false positive is much more expensive than a miss, then change the data enough to see when its confidence falls apart. If it stays confident through that, the benchmark number doesn't mean much.
Cloudflare also announced a fine-tuning service that lets customers adapt Clef to their own workloads using their data, initially with support from Cloudflare engineers, with a self-service platform planned for a later release. No firm date has been announced.
The company has made the Clef models available through Workers AI and as downloadable weights on Hugging Face, inviting developers to experiment with them and provide feedback.
About the Author
Renato Losio
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

