Alberta industrial teams collect photos, recordings and short video clips that someone later has to look at and sort: is the label visible, does the fan sound right, is the valve leaking? A decision model doesn’t describe what it sees. It answers the specific questions you ask, with a probability for each answer. Cloudflare’s new Clef-omni does that across four kinds of input in one call.
What Cloudflare released
Cloudflare’s post says Clef-omni takes audio and video input alongside text and images, and that decision models had been mostly text-only until now. The post also says it cut the price of Clef-flash and made Clef faster. Clef-omni is built on Alibaba’s Qwen3-Omni-30B-A3B-Instruct mixture-of-experts model, with the text-to-speech output parts removed. Cloudflare says it scores all modalities and answer options in a single pass instead of generating tokens.
“Instead of setting up cascading pipelines of models that transcribe speech-to-text, or splitting audio and image channels from video, you can just call one model to make decisions across any modality.”Cloudflare, Introducing Clef-omni, October 9, 2026
The speed figures are Cloudflare’s own: text-only decisions in about 130 ms at the median, image inputs in about 150 ms, and a 21-second video clip with sound in about 1.5 seconds. Cloudflare’s benchmark tables are also its own and are mixed: Clef-omni scores higher than Clef on some tests and lower on others, so test it on your data. The base Clef and Clef-flash models were launched on the same Cloudflare blog on October 1, in a post titled Introducing Clef; we have not covered that launch before.
Prices, context and limits
| Model | Input price | Hosted context |
|---|---|---|
| Clef-flash | $0.038 (was $0.09) | 24K (was advertised as 64K) |
| Clef | $0.24 | 64K |
| Clef-omni | $0.15 | 64K |
Cloudflare’s model page lists the request limits: up to 4 images, 4 audio clips (8 MiB and 300 seconds each) and 2 videos (16 MiB and 60 seconds each), with audio and video together capped at 16 MiB. Remote URLs aren’t accepted, so media must be sent embedded. Media is converted to input tokens and billed at the input rate: the page lists about 780 tokens per minute of audio, and up to 1,024 tokens per image. Confirm the billing currency and taxes on your own account.
The Hugging Face card says the weights are released under Apache-2.0, following the base model, and that the backbone needs about 64 GB of GPU memory in bfloat16. That gives you a self-hosting option if hosted processing doesn’t fit your rules. Neither Cloudflare’s post nor the model page states which regions process hosted Workers AI requests, so we don’t print one.
Where it could fit in Alberta (our advice)
What follows is our advice, not Cloudflare’s. Two uses stand out. The first is inspection triage: send a site photo, a short recording and a clip, then ask fixed questions such as “is the equipment label legible?” or “does the motor sound normal?” and have a technician review anything below a confidence threshold. Cloudflare’s own example asks nearly the same questions about a unit installation. The second is intake classification for forms, voicemails and photos sent to a service business.
Pilot it on non-sensitive material first, such as public or staged photos, and check where processing happens before sending anything about a customer, a site layout or a safety system. If the answer isn’t clear, ask Cloudflare in writing, or consider the open weights on your own hardware. Treat scores as triage hints, never as sign-off on safety or compliance. Our industrial AI Alberta page covers plant and field use cases, private AI security covers the data boundary, and an AI consulting engagement can design the pilot.
