cloudflare3 min read

Curated summary

Powering the agents: Workers AI now runs large models, starting with Kimi K2.5

Read original(opens in new tab)

Cloudflare is expanding Workers AI beyond smaller models by adding Moonshot AI’s Kimi K2.5, a frontier open-source model designed for agentic workloads. With a 256k context window, tool calling, vision, and structured outputs, Kimi can power an agent’s full lifecycle directly on Cloudflare’s platform. Cloudflare argues that its price-performance makes open-source models essential as personal and enterprise agents dramatically increase inference demand.

Kimi K2.5’s Price-Performance Advantage

  • Cloudflare uses Kimi internally for:
    • Agentic coding through OpenCode
    • Automated code review via the Bonk public code review agent
    • Security analysis of Cloudflare codebases
  • A security-review agent processes more than 7 billion tokens daily and has found over 15 confirmed issues in one codebase.
  • Compared with a mid-tier proprietary model, switching to Kimi reduced the estimated cost of this workload by 77%, from roughly $2.4 million annually.
  • As employees increasingly run multiple agents continuously, inference costs become a major barrier to scaling.
  • Cloudflare positions open-source, frontier-quality models as a more economical alternative to proprietary systems.

Serving Large Models on Workers AI

  • Supporting Kimi required upgrades to Workers AI’s inference stack, which historically focused on smaller models.
  • Cloudflare uses its proprietary Infire inference engine and custom kernels to improve:
    • Model performance
    • GPU utilization
    • Throughput
  • The platform applies advanced serving strategies such as:
    • Data, tensor, and expert parallelization
    • Disaggregated prefill, separating input processing from generation across machines
  • Workers AI handles these infrastructure optimizations so developers do not need specialized machine learning, DevOps, or reliability engineering expertise.

Prefix Caching for Agent Workloads

  • Agents frequently resend large prompts containing:
    • System instructions
    • Tool definitions
    • MCP server tools
    • Conversation history
    • Entire codebases
  • Prefix caching avoids reprocessing unchanged input tokens during multi-turn interactions.
  • This reduces prefill work, improving:
    • Time to First Token (TTFT)
    • Tokens Per Second (TPS)
    • Overall inference cost
  • Workers AI now exposes cached tokens as a usage metric and charges less for them than regular input tokens.
  • Cloudflare has also introduced techniques to improve cache hit rates.

Session Affinity

  • Workers AI provides an x-session-affinity header to route requests from the same session or agent to the same model instance.
  • Keeping requests on the same instance increases prefix-cache reuse.
  • Higher cache hit rates lead to faster responses, greater throughput, and lower costs.
  • Clients should provide a unique session or agent identifier with the header.

Cloudflare’s recommendation is to use Workers AI when building agents that need frontier-level reasoning without the cost and operational burden of proprietary models or self-hosted infrastructure.

Continue with another curated summary.