Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

Securing non-human identities: automated revocation, OAuth, and scoped permissions

Cloudflare argues that securing modern infrastructure requires managing non-human identities—agents, scripts, and third-party applications—as carefully as human users. The core model combines principals, credentials, and policies, with protections covering token leakage, OAuth access visibility, and narrowly scoped permissions. The post focuses especially on automated token detection and revocation, designed to limit damage when credentials are exposed. ## Identity as Three Connected Components - **Principal:** The identity acting on a system’s behalf, such as a developer, AI agent, background service, or OAuth application. - **Credential:** The proof of identity, typically an API token. Anyone who obtains it may impersonate the principal. - **Policy:** The permissions assigned to the identity, determining which resources and actions it can access. - Security failures occur when these elements are managed separately—for example, when a valid identity uses a stolen token or has unnecessarily broad permissions. ## Automated Detection and Revocation of Leaked Tokens - API tokens are commonly exposed by accidentally committing them to public repositories. - Cloudflare cites GitGuardian’s estimate that more than 28 million secrets were published to public GitHub repositories in the previous year, with AI-driven development increasing leak rates. - Cloudflare is partnering with credential-scanning providers to detect leaked tokens and revoke them before attackers can exploit them. - New Cloudflare token formats use a recognizable `cf` prefix and a checksum, allowing scanners to identify tokens confidently and verify whether they are authentic. - Existing tokens remain valid, but newly generated tokens use the scannable format. ## GitHub Secret Scanning Integration - GitHub scans public and private repositories for the new Cloudflare token formats on every commit. - For public repository leaks: - GitHub validates the token using its checksum. - GitHub sends Cloudflare a webhook. - Cloudflare automatically revokes the token. - The user receives an email prompting them to create a replacement. - For private repositories, GitHub notifies the customer so the leaked credential can be removed and replaced. ## Protection Through Cloudflare One Cloudflare One customers can use the Credentials and Secrets DLP profile to detect and block Cloudflare tokens across multiple data paths: - **Network traffic:** Cloudflare Gateway can block tokens in uploads, downloads, and outbound requests. - **Email:** Cloudflare Email Security and the DLP Assist add-in can scan Microsoft 365 messages before external delivery. - **Stored data:** Cloudflare CASB scans connected services such as Google Drive, OneDrive, and Dropbox. - **AI traffic:** Cloudflare AI Gateway can inspect prompts and model responses in real time, addressing credential exposure through AI systems. ## Broader Scanner Ecosystem - Cloudflare is working with open-source and commercial credential-scanning tools. - The goal is to protect customers regardless of which repository or secret-scanning products they use. - Automatic revocation is presented as a critical safeguard because credential exposure is treated as inevitable rather than exceptional. Organizations should use recognizable, verifiable tokens, enable repository and DLP scanning, and ensure leaked credentials are revoked automatically. These controls reduce the window in which an exposed token can be used and should be combined with narrowly scoped permissions for agents and applications.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Scaling MCP adoption: Our reference architecture for simpler, safer and cheaper enterprise deployments of MCP

Cloudflare argues that enterprise MCP adoption requires centralized governance rather than individually managed, locally hosted servers. Its reference architecture combines remote MCP servers, Cloudflare Access, MCP server portals, and AI security controls to improve visibility, authentication, policy enforcement, and performance. The company also introduces Code Mode with MCP server portals to reduce the token and context-window costs of exposing large APIs. ## Centralized Remote MCP Servers - MCP separates the AI application from corporate credentials and APIs: - The MCP client connects to the LLM or agent. - The MCP server mediates access to internal resources. - Cloudflare moved away from locally hosted MCP servers because they: - May use unvetted software and versions. - Increase supply-chain and tool-injection risks. - Are difficult for IT and security teams to administer. - A centralized team manages MCP infrastructure through a shared monorepo platform. - Approved teams can create governed MCP servers from templates, inheriting: - Default-deny write controls. - Audit logging. - Automated CI/CD pipelines. - Secrets management. - Servers are deployed remotely on Cloudflare’s developer platform and custom domains, providing centralized usage visibility and global low-latency access. ## Authentication with Cloudflare Access - Public MCP servers, such as documentation and Radar services, can remain openly accessible. - MCP servers connected to private corporate resources require employee authentication. - Cloudflare Access acts as the OAuth provider and identity layer. - It verifies: - Single sign-on. - Multifactor authentication. - IP address, location, and device-certificate context. - Access issues tokens that authorize users to reach protected resources. ## MCP Server Portals for Discovery and Governance - As the number of MCP servers grew, employees needed a central way to discover authorized services. - Users connect their MCP client to a portal, which exposes the internal and third-party MCP servers they are permitted to use. - Portals provide: - Centralized logging. - Consistent policy enforcement. - Data loss prevention controls. - Access policies for users and tools. - Administrators can restrict both portal access and the specific tools exposed by each server. - Finance users might receive only read-only repository tools. - Engineering users on corporate devices might receive read/write capabilities. - Portals support MCP servers hosted on Cloudflare as well as third-party servers. - Cloudflare emphasizes that the relevant security and networking components can run on the same physical machine in its global network, reducing latency and avoiding unnecessary traffic transit. ## Code Mode Reduces MCP Token Costs - The standard MCP design exposes every API operation as a separate tool. - For large platforms with thousands of endpoints, this exhaustive tool list consumes an agent’s context window and increases token costs. - Cloudflare presents Code Mode with MCP server portals as a way to address this scaling problem. - The provided article excerpt ends while introducing Cloudflare’s earlier use of server-side Code Mode for exposing large numbers of API endpoints. Cloudflare’s approach recommends treating MCP as enterprise infrastructure: centrally deployed, authenticated, discoverable, policy-controlled, and monitored. Organizations adopting MCP at scale should avoid unmanaged local servers and provide reusable platforms that make secure deployment the default.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Managed OAuth for Access: make internal apps agent-ready in one click

Cloudflare’s managed OAuth makes internal apps behind Cloudflare Access usable by AI agents without modifying the apps themselves. By exposing standard OAuth discovery and authorization flows, agents can authenticate on behalf of the human user rather than relying on static service accounts. The result is immediate agent compatibility for legacy websites, APIs, and other internal tools. ## The Problem: Access Worked for Humans, Not Agents - Cloudflare protects thousands of internal and self-hosted applications with Cloudflare Access. - Humans can follow Access’s login-page redirect, but agents generally cannot interact with browser-based authentication flows. - Cloudflare initially addressed this internally by modifying OpenCode’s web fetch tool to use `cloudflared` to obtain a JWT and attach it to requests. ## Managed OAuth for Access Applications - Managed OAuth is now available in open beta for every Access application. - Enabling it requires one click and no application code changes. - Access acts as the OAuth authorization server and advertises authentication details through: - The `WWW-Authenticate` response header - `/.well-known/oauth-authorization-server` - OAuth-capable agents can then: - Dynamically register as clients using RFC 7591. - Send the user through a PKCE authorization flow using RFC 7636. - Receive a token representing the user’s authorization. - The same pattern supports web pages, web applications, REST APIs, and MCP servers. ## Making Legacy Internal Apps Agent-Ready - Retrofitting every internal application with APIs, CLIs, MCP servers, and new agent standards is impractical. - Many applications can already provide useful value when agents treat them as ordinary websites. - For example, an internal wiki may only need Markdown-for-Agents support and managed OAuth. - Putting Cloudflare Access in front of existing applications provides immediate agent compatibility without rebuilding them. ## User-Based Authorization Instead of Service Accounts - Static service accounts and tokens can be useful for simple integrations, but they weaken attribution and fine-grained access control. - Actions performed through shared credentials may appear in audit logs as originating from the agent or service account rather than the responsible human. - They can also create confused-deputy risks, where an agent gains authority beyond what its user should have. - OAuth preserves the user–agent relationship: - Tokens are scoped to the user’s identity and permissions. - Existing access policies continue to apply. - Audit logs can attribute actions to the initiating user. ## RFC 9728 and Agent Web Fetching - RFC 9728 standardizes how clients discover OAuth authentication requirements. - MCP has adopted the standard, but Cloudflare argues that general-purpose agents should use it for protected websites and REST APIs as well. - Most agent web-fetch tools currently ignore `WWW-Authenticate` headers and do not automatically: - Locate the OAuth authorization-server metadata. - Register as an OAuth client. - Complete the authorization flow. - Cloudflare has drafted changes to OpenCode’s web-fetch tool demonstrating how tools could check for existing credentials and initiate OAuth when necessary. Cloudflare’s recommendation is to enable managed OAuth for Access-protected applications and encourage agent developers to implement RFC 9728. This offers a practical path to agent adoption while retaining user-level permissions, accountability, and compatibility with existing internal software.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Secure private networking for everyone: users, nodes, agents, Workers — introducing Cloudflare Mesh

Cloudflare introduces Mesh as a private networking layer designed for humans, services, and autonomous AI agents. It connects devices, servers, cloud VPCs, Workers, Durable Objects, and Agents SDK applications without exposing private services publicly or relying on manual VPN and SSH workflows. Mesh builds on Cloudflare One, so existing Gateway policies, Access rules, device posture checks, and other Zero Trust controls apply automatically. ## Why Agent Workloads Need Private Networking - AI agents increasingly need access to private databases, APIs, repositories, MCP servers, object stores, and home infrastructure. - Traditional solutions are poorly suited to autonomous software: - VPNs often require interactive login. - SSH tunnels require manual setup. - Public exposure increases the risk of unauthorized access. - Basic connectivity does not provide sufficient visibility into agent activity. - Agents may have powerful permissions, including shell, filesystem, and network access, making misconfiguration especially dangerous. ## New Agentic Workflows - **Accessing personal agents remotely** - A user can run an agent such as OpenClaw on a home Mac mini. - Phones, laptops, and work devices can connect securely without exposing the agent directly to the public Internet. - **Letting coding agents access staging systems** - Agents such as Claude Code, Cursor, or Codex can reach private staging databases, analytics systems, APIs, and object stores. - Developers avoid exposing those systems or tunneling an entire laptop into a cloud VPC. - **Connecting deployed agents to private services** - Agents running on Cloudflare Workers can call internal APIs and databases. - Mesh is intended to provide scoped access, auditability, and reduced credential exposure. ## How Cloudflare Mesh Works - Mesh uses a lightweight connector and a single binary to connect: - Personal devices - Remote servers - User endpoints - Private cloud networks - Connected devices communicate over private IPs through Cloudflare’s global network, which spans more than 330 cities. - Cloudflare’s existing terminology is simplified: - WARP Connector becomes a **Cloudflare Mesh node**. - WARP Client becomes the **Cloudflare One Client**. - Example deployments include: - Connecting an iPhone to a home Mac mini running an agent. - Connecting a developer laptop to staging databases and internal APIs. - Connecting Linux servers and external cloud VPCs so agents can reach private resources and MCP servers. ## Security and Cloudflare One Integration - Mesh traffic automatically inherits Cloudflare One protections, including: - Gateway network, DNS, and HTTP policies - Device posture checks - DNS filtering - Access rules - Existing Cloudflare One customers can use Mesh without adopting a separate security platform. - Organizations can later expand into: - Access for Infrastructure for SSH and RDP management - Browser Isolation - Data Loss Prevention - Cloud Access Security Broker capabilities - The goal is to protect agent traffic with the same controls already used for human users and services. Cloudflare Mesh is positioned as a practical starting point for securely connecting agents to private infrastructure. Teams can begin with simple private networking and later add more advanced Zero Trust controls without migrating to a different platform.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab and Vertex AI on Google Cloud: Advancing agentic development

GitLab is partnering with Google Cloud to combine the GitLab Duo Agent Platform’s lifecycle-wide orchestration with Vertex AI’s managed foundation models and enterprise controls. The integration gives development teams context-aware agents for planning, coding, security, and delivery while keeping workflows within GitLab’s governed system of record. Customers gain model flexibility, stronger governance, and reduced complexity compared with managing disconnected AI tools. ## Agents Across the Software Development Lifecycle - GitLab Duo Agent Platform coordinates specialized agents across planning, development, code review, security, and delivery. - Unlike standalone coding assistants, GitLab agents can access issues, merge requests, pipelines, vulnerabilities, and codebases. - GitLab Duo Planner Agent can analyze backlogs, divide epics into tasks, and support prioritization. - Security Analyst Agent can triage vulnerabilities, explain risks, and recommend remediation priorities. - Built-in flows connect agents into end-to-end processes, reducing manual handoffs. - Agentic Chat provides natural-language access to project context and multi-step reasoning within GitLab. ## Vertex AI as the Model and Infrastructure Layer - Vertex AI supplies the foundation models and related services used by GitLab agents. - Newer models improve reasoning, tool use, and long-context understanding, supporting workloads such as backlog analysis and monorepo security reviews. - Vertex AI Model Garden offers Gemini, third-party, and open-source models, allowing customers to balance performance, cost, and regulatory requirements. - GitLab supports Bring Your Own Model configurations, enabling organizations to use approved providers and gateways. - Vertex AI abstracts LLM hosting, including infrastructure management, security, governance, and model-version delivery. ## Enterprise Governance and Operational Benefits - GitLab’s AI Gateway mediates model access, helping administrators track connections and maintain governance. - Developers remain in GitLab while inference follows existing Google Cloud security and policy controls. - Platform teams can standardize which models support recommendations, analysis, and remediation. - Security teams can manage findings and proposed fixes in the same environment, reducing context switching and unmanaged workflows. - Using Vertex AI through GitLab can align AI usage with existing Google Cloud contracts, controls, and procurement policies. - The approach helps reduce duplicate spending and fragmented “shadow AI” toolchains. ## Practical Outcome for Google Cloud Customers The integration is intended to increase developer productivity without requiring teams to evaluate, host, or manage individual language models. GitLab provides the governed DevSecOps control plane, while Vertex AI supplies scalable, flexible model infrastructure, enabling organizations to adopt more capable agentic workflows while maintaining enterprise security and control.

Read original(opens in new tab)
slack4 min readCurated summary

Managing context in long-run agentic applications

Long-running multi-agent applications cannot rely on unlimited conversation history: model APIs are stateless, and growing context windows eventually reduce quality or hit hard limits. Slack’s security-investigation system addresses this by giving agents complementary, purpose-specific context rather than exposing every agent to the full investigation history. Its three main channels—the Director’s Journal, Critic’s Review, and Critic’s Timeline—preserve coherence while leaving room for independent reasoning. ## The Challenge of Long-Run Coherence - Agent frameworks usually maintain continuity by resending the complete message history with every inference request. - Long investigations can involve hundreds of requests and megabytes of generated output. - Context windows impose both: - A hard limit on how much history can be supplied. - A quality limit, because performance may degrade before the window is completely full. - Multi-agent systems need carefully scoped views: - Too little shared context makes agents disconnected from the investigation. - Too much shared context can suppress creativity and encourage confirmation bias. ## Three Complementary Context Channels Slack uses separate information sources for different purposes: - **Director’s Journal** - Structured working memory for the orchestrating Director. - Records decisions, observations, findings, questions, actions, and hypotheses. - **Critic’s Review** - An annotated report evaluating Expert findings. - Includes credibility scores to distinguish reliable evidence from weaker claims. - **Critic’s Timeline** - A consolidated chronological view of findings. - Also attaches credibility scores, helping agents understand the sequence and evidential strength of events. Together, these channels provide continuity without forcing every agent to process the entire raw conversation. ## The Director’s Journal The Director coordinates the investigation by choosing questions, assigning specialist Experts, assessing progress, and deciding when to stop. The Journal gives it persistent working memory across phases and rounds. - The Director is encouraged to update the Journal frequently with short notes. - Entries can represent: - **Decisions** about investigative strategy - **Observations** about emerging patterns - **Findings** representing confirmed facts - **Questions** that remain unresolved - **Actions** taken or planned - **Hypotheses** about what may be happening - Entries can also include: - Priority levels - Follow-up actions - References to supporting evidence - Investigation phase, round number, and timestamp - The journaling tool itself simply accumulates entries; the agents’ prompts explain how to interpret them. ## Maintaining Alignment Across Agents - The Journal creates a shared narrative around the Director’s evolving plan. - It helps the Director: - Track progress - Identify dead ends - Revise investigative direction - Preserve decisions between rounds - Guide other agents toward a conclusion - Every agent receives the current Journal chronologically, along with instructions describing: - The Director’s role - Each agent’s relationship to the Director - The Journal’s purpose - How its entries should influence their work - This approach keeps specialists anchored to the overall investigation without requiring them to read every prior interaction. ## Example Investigation Context The sample Journal comes from an investigation into an apparent kernel-module-loading alert that turned out to be a false positive. - The Director recorded that: - The event originated from a package-installation hook rather than a direct `modprobe` command. - The host appeared to be a personal development workstation. - Root access was expected in that environment. - The detection rule matched “kmod” in a script path rather than confirming module loading. - The Director identified relevant Expert domains, including: - Endpoint telemetry - Identity and access - Configuration management - User behavior - The Journal captured both the preliminary conclusion and remaining verification tasks, such as checking the parent process chain. The design therefore preserves the reasoning trail while keeping it structured and compact. A practical design for long-running agentic systems is to replace indiscriminate transcript accumulation with multiple, curated context channels. Persistent journals can maintain leadership and continuity, while independent reviews and timelines provide evidence-focused context without overwhelming agents or biasing their reasoning.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: Claude Mythos Preview in Amazon Bedrock, AWS Agent Registry, and more (April 13, 2026) | Amazon Web Services

AWS’s April 13, 2026 roundup centers on improving governance and visibility as organizations move AI workloads into production. Amazon Bedrock added IAM user and role-based cost allocation, while Claude Mythos Preview and the AWS Agent Registry expanded capabilities for cybersecurity and agent management. The week also brought updates across storage, observability, WorkSpaces, and quantum computing. ## Bedrock Cost Allocation - Organizations can tag IAM users and roles with attributes such as team or cost center. - Activated tags appear in Billing and Cost Management, AWS Cost Explorer, and detailed Cost and Usage Reports. - This enables teams to track foundation model inference costs across departments, agents, and tools such as Claude Code on Bedrock. ## Claude Mythos Preview in Amazon Bedrock - Anthropic’s Claude Mythos is available as a gated research preview through Project Glasswing. - The model is designed for advanced cybersecurity work, including: - Finding sophisticated vulnerabilities - Analyzing large codebases - Handling complex reasoning and coding tasks - Access is limited to allowlisted organizations, with priority given to critical internet companies and open-source maintainers. ## AWS Agent Registry - AgentCore’s new registry provides a private catalog for AI agents, tools, skills, MCP servers, and custom resources. - Features include semantic and keyword search, approval workflows, and CloudTrail auditing. - Teams can access it through the AgentCore Console, AWS CLI, SDKs, or as an MCP server from IDEs. - The goal is to improve reuse and governance instead of having teams independently recreate capabilities. ## Other AWS Launches - **Amazon S3 Files:** Exposes S3 buckets as shared file systems with file-system semantics, caching, and high aggregate read throughput. Applications can use file-system and S3 APIs simultaneously without migration or code changes. - **OpenSearch observability:** Adds Managed Prometheus, PromQL support, RED metrics, agent tracing, and OpenTelemetry GenAI semantic conventions for correlating AI execution with logs and traces. - **WorkSpaces Advisor:** Uses generative AI to diagnose Amazon WorkSpaces Personal configuration issues and recommend fixes. - **Amazon Braket:** Adds Rigetti’s 108-qubit Cepheus-1-108Q processor, supporting Braket SDK, Qiskit, CUDA-Q, Pennylane, and pulse-level control. ## Additional Resources and Upcoming Events - AWS highlighted guidance for regional availability monitoring with S3, Bedrock model lifecycle management, memory-intensive Lambda managed instances, and OpenClaw deployment choices. - Kiro is bringing back startup credits, offering eligible companies one year of Pro+ access across three team-size tiers. - The virtual “What’s Next with AWS” event on April 28 will focus on agentic AI and feature AWS, OpenAI, and industry leaders. Organizations adopting AI at scale should prioritize IAM-based cost attribution, centralized agent governance, and lifecycle planning for foundation models.

Read original(opens in new tab)
github1 min readCurated summary

GitHub for Beginners: Getting started with GitHub Pages

Kedasha is a GitHub Developer Advocate who shares her software development experience and industry knowledge with the broader developer community. She is passionate about helping others learn and can be found online as **@itsthatladydev**. ### Role and Focus - Works as a Developer Advocate at GitHub. - Shares lessons from her professional experience with developers. - Helps others learn about the technology industry. ### Online Presence - Identifies online as **@itsthatladydev**.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Building a CLI for all of Cloudflare

Cloudflare is rebuilding Wrangler into a unified CLI for its entire platform, motivated by the growing role of coding agents in configuring and deploying Cloudflare applications. The technical preview, available as `npx cf` or the globally installed `cf` package, currently covers only a subset of products but is intended to support the full API surface. The effort depends on a new TypeScript-based schema system that can generate consistent commands, configuration, bindings, documentation, and agent-oriented interfaces. ## A CLI for all of Cloudflare - Cloudflare offers more than 100 products and nearly 3,000 HTTP API operations. - Agents increasingly use Cloudflare APIs to: - Build and deploy applications - Configure accounts - Query analytics and logs - Create agents and platforms - Cloudflare aims to expose its products consistently through: - CLI commands - Workers Bindings - SDKs - Configuration files - Terraform - Documentation and OpenAPI schemas - MCP servers and Agent Skills - The new Wrangler technical preview can be tried with: - `npx cf` - `npm install -g cf` - A broader internal version already supports the full Cloudflare API, with ongoing work to make command output useful for both humans and agents. ## A new schema and code-generation pipeline - Existing OpenAPI schemas already generate: - Cloudflare SDKs - The Terraform provider - The Code Mode MCP server - Other interfaces, including Wrangler commands, Workers Bindings, configuration, documentation, and Agent Skills, were previously maintained manually. - Manual synchronization was error-prone and could not scale to Cloudflare’s full product range. - OpenAPI alone is insufficient because it primarily describes REST APIs, while Cloudflare also needs to represent: - Interactive CLI workflows - Multiple local and remote actions - RPC-style Workers Bindings - Agent Skills and related documentation - Cloudflare therefore created a TypeScript schema format containing: - API definitions - CLI commands and arguments - Context required to generate different interfaces - Conventions, linting, and guardrails enforce consistency while allowing the schema to generate OpenAPI and future interfaces. ## Consistency for agents and humans - Agents depend on predictable command names and flags. Inconsistent syntax can cause them to call commands that do not exist. - Cloudflare is enforcing conventions at the schema layer, including: - `get`, never `info` - `--force`, never `--skip-confirmations` - `--json`, never `--format` - Applying these rules across interfaces avoids discrepancies between the CLI, REST APIs, and SDKs. - Wrangler must also clearly distinguish local and remote resources. - This is especially important for D1, R2, and KV, where local simulation and remote bindings can coexist. - Clear defaults and output indicating whether an operation targets local or remote resources help agents avoid modifying the wrong environment. ## Local Explorer for simulated resources - Local Explorer is available in open beta through Wrangler and the Cloudflare Vite plugin. - It lets developers inspect locally simulated: - KV - R2 - D1 - Durable Objects - Workflows - Local resources use the same underlying API structure as Cloudflare’s remote APIs and Dashboard. - Cloudflare’s local development environment runs Workers APIs locally, including D1 backed by SQLite through Miniflare. - Previously, developers had to inspect `.wrangler/state` or use third-party tools to understand local data. - Local Explorer provides an interface showing: - Which bindings are attached to a Worker - What data those bindings contain - It can be opened with the `e` keyboard shortcut and helps developers or agents verify schemas, seed test data, and reset local databases. Cloudflare’s direction is to make Wrangler a consistent, machine-readable interface to the entire platform. The technical preview is early, but the new schema-driven system and Local Explorer establish the foundation for a CLI that is easier for both developers and coding agents to use safely.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Durable Objects in Dynamic Workers: Give each AI-generated app its own database

Dynamic Workers make it possible to run AI-generated code securely in lightweight isolates, but disposable execution is not enough for persistent applications. Cloudflare’s Durable Object Facets address this by letting a supervised Durable Object dynamically load an AI-generated Durable Object class with its own SQLite-backed storage. This combines sandboxed, persistent application state with centralized control over provisioning, access, logging, metrics, and billing. ## From Disposable Code to Persistent Apps - Dynamic Workers load code on demand in secure isolates rather than containers. - Isolates start quickly and use little memory, making them suitable for short-lived AI-generated tasks. - Persistent AI-built applications need: - Custom user interfaces - Long-lived state - Secure execution - A remote SQL database could provide storage, but it introduces network latency and additional infrastructure. ## Why Durable Objects Fit - Each Durable Object has: - A globally unique name - One active instance per name - An attached SQLite database stored locally - Local SQLite storage provides extremely low-latency access. - AI-generated applications can therefore use normal Durable Object storage APIs, including key-value and SQL storage. ## Limitations of the Traditional Model - Standard Durable Objects require: - A class extending `DurableObject` - Exporting the class from the Worker - Wrangler configuration to provision storage - A namespace binding for access - This model does not naturally support code loaded dynamically at runtime. - Giving an agent direct control of Durable Object namespaces could also allow uncontrolled object creation and storage use. - A platform needs an intermediary to enforce limits and provide observability, billing, and other operational controls. ## Durable Object Facets - Facets allow a normal, statically configured Durable Object to dynamically instantiate another Durable Object class. - The outer object acts as a supervisor: - Loads the agent’s code as a Dynamic Worker - Selects the exported Durable Object class - Forwards requests or RPC calls - Controls and monitors the application - The dynamically loaded class can directly extend `DurableObject`. - Each facet receives its own SQLite database, separate from the supervisor’s database. - Multiple facets can exist within one Durable Object, each identified by a name and subject to storage limits. ## Example Architecture - An `AppRunner` Durable Object receives incoming requests. - It obtains a facet named `"app"` through `this.ctx.facets.get(...)`. - When the facet starts, the runner: - Loads the Dynamic Worker - Retrieves its exported application class - Instantiates it as the facet - Requests are then forwarded to the dynamically loaded application. - The sample application maintains a request counter using Durable Object storage. Durable Object Facets provide a practical foundation for AI-generated applications that need persistent state without sacrificing isolation or platform governance. They are especially suited to personal or small “vibe-coded” apps, where each application can receive its own storage while the host platform retains control over resource usage and operational policies.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Agents have their own computers with Sandboxes GA

Cloudflare has made Sandboxes and Cloudflare Containers generally available for running AI agents in persistent, isolated computer environments. The platform addresses the operational challenges of agent workloads—including bursty demand, fast state restoration, secure authentication, lifecycle control, and developer-friendly tooling. Recent additions make Sandboxes more capable for coding agents while reducing the cost of running them at scale. ## Why Agents Need Full Computers - Coding agents often need to clone repositories, build software, run development servers, and work across multiple languages. - Existing VM and container approaches must handle: - Rapidly creating many session-specific environments without paying for idle capacity. - Quickly restoring previous session state. - Giving agents access to services without exposing credentials. - Programmatic control over commands, files, and sandbox lifecycles. - Simple interfaces for both human developers and agents. - Figma is using Cloudflare Containers to run untrusted agent- and user-authored code for Figma Make. ## Sandboxes 101 - A Sandbox is a persistent, isolated environment powered by Cloudflare Containers. - Sandboxes are addressed by name: - Running sandboxes are reused. - Inactive sandboxes sleep automatically. - Requests wake sleeping sandboxes on demand. - The same sandbox can be accessed from anywhere using its ID. - The API supports operations such as: - `exec` for running commands. - `gitCheckout` or `gitClone` for retrieving repositories. - `writeFile` for managing files. - Command output can be streamed in real time, such as when running `npm test`. ## Secure Credential Injection - Agents may need to call private services but should not receive raw credentials. - Sandboxes inject credentials at the network layer through a programmable egress proxy. - Custom outbound rules can add authentication headers to requests based on the destination host. - This allows authenticated access while keeping secrets outside the agent’s environment. - Authentication logic can be customized for identity-aware access, dynamic rules, and Workers bindings. ## Real Terminal Access with PTY - Early agent interfaces treated shell commands as isolated request-response operations. - PTY support provides a more realistic terminal experience: - Output streams continuously. - Processes can be interrupted. - Sessions can be reconnected later. - Sandbox terminal sessions are proxied over WebSockets and are compatible with `xterm.js`. - Applications can expose the backend through `sandbox.terminal`. ## Features for Agent Development - **Persistent code interpreters:** Stateful Python, JavaScript, and TypeScript execution is available out of the box. - **Background processes:** Development servers and other long-running commands can continue running independently. - **Live preview URLs:** Agents and users can inspect development servers and verify changes while they are in progress. - **Filesystem watching:** Faster feedback as agents modify files. - **Snapshots:** Coding sessions can be quickly recovered from saved state. - **Higher limits and Active CPU Pricing:** Fleets of agents can scale without paying for unused CPU cycles. Cloudflare’s GA release positions Sandboxes as a managed environment for agent-driven software development: persistent when state matters, isolated when code is untrusted, and cost-efficient when workloads are intermittent.

Read original(opens in new tab)
cloudflare4 min readCurated summary

Dynamic, identity-aware, and secure Sandbox auth

Sandboxes for AI agents need more than isolation: they also require fast startup, platform control, and safe access to external services. The post introduces outbound Workers, programmable egress proxies that intercept sandbox traffic and can authenticate, restrict, modify, log, or cancel requests. This approach combines zero-trust security with identity-aware, flexible, observable, and dynamic authorization without exposing secrets to untrusted agents. ## Sandbox Requirements Sandboxes provide three core benefits: - **Security:** Untrusted users or agents can run code without compromising the host or neighboring sandboxes, often through microVM isolation. - **Speed:** Users can quickly start new sandboxes and restore existing state. - **Control:** The trusted platform can mount files, execute commands, and control network access inside the sandbox. Outbound Workers add network-level control to this model by acting as programmatic egress proxies for Sandboxes and Containers. ## How Outbound Workers Work - A sandbox can define handlers for all outbound requests or for requests to specific hosts. - For example, requests to `github.com` can be intercepted through `static outboundByHost`. - The handler can: - Add authentication headers. - Log requests. - Modify request data. - Reject or cancel requests. - Secrets remain outside the sandbox and can be accessed by the Worker through its environment. - Workers run near the sandbox, can access distributed state, and can be updated using ordinary JavaScript. A sample handler copies the request headers and injects `x-auth-token` from `env.SECRET` before forwarding the request. ## Challenges with Existing Agent Authentication Agent workloads cannot be fully trusted, even when the underlying language model is not intentionally malicious. Credentials must therefore limit accidental misuse and prevent data exfiltration. ### Standard API Tokens - Tokens are commonly passed through environment variables or mounted secret files. - They are simple to implement but expose credentials to the sandboxed workload. - A compromised or misbehaving agent could leak the token. - Expiration and rotation are required, creating operational overhead. ### Workload Identity Tokens - Systems such as OIDC provide an identity assertion rather than a general-purpose service token. - The agent can exchange the identity token for a short-lived access token. - Tokens can be invalidated when a workflow ends, simplifying expiration. - The drawback is limited upstream support: many services do not natively accept OIDC, forcing platforms to build custom token-exchange services. ### Custom Proxies - Proxies provide maximum control and can enforce granular permissions even when an upstream service has weak RBAC. - They can be combined with workload identity tokens. - However, intercepting all sandbox traffic and building an efficient, dynamic, programmable proxy is difficult. ## Characteristics of an Ideal Agent Auth System The post argues that agent authentication should be: - **Zero trust:** Never expose a reusable token to an untrusted workload. - **Simple:** Avoid complicated token minting, rotation, and decryption systems. - **Flexible:** Enforce permissions independently of the upstream service. - **Identity-aware:** Apply rules based on which sandbox is making the request. - **Observable:** Record and inspect outbound calls. - **Performant:** Avoid slow, centralized authorization round trips. - **Transparent:** Require no changes to the sandboxed application. - **Dynamic:** Allow authorization rules to change while systems are running. Outbound Workers are presented as a way to satisfy all of these requirements. ## Restriction and Observability A basic outbound handler can enforce network policy with only a few lines of JavaScript: - Inspect each outgoing HTTP request. - Log requests using disallowed methods. - Return a `405 Method Not Allowed` response for anything other than `GET`. - Forward permitted requests with `fetch(req)`. This demonstrates that outbound Workers can enforce restrictions and provide observability without modifying the application running inside the sandbox. ## Practical Recommendation Use outbound Workers as a trusted egress layer for agent sandboxes. Keep sensitive credentials outside the workload, inject or exchange them only at the proxy, and use the Worker to enforce identity-specific policies, logging, and request restrictions dynamically.

Read original(opens in new tab)
grammarly3 min readCurated summary

The Trust Question: How Higher Education Is Really Navigating AI

Higher education’s AI challenge is fundamentally a question of trust, not simply technology adoption or resistance. Based on interviews with educators and administrators, institutions are balancing competing views about innovation, evidence, ethics, and practical outcomes. The central conclusion is that credible AI governance must make these differences visible and build shared understanding rather than rely on blanket rules or binary narratives. ## AI Policy Is a Campus-Wide Negotiation - Four recurring orientations shape institutional responses: - **Innovators** favor responsible adoption before reactive governance becomes necessary. - **Strategists** want stronger evidence before committing to change. - **Resisters** prioritize ethics, academic integrity, and institutional reputation. - **Pragmatists** focus on student success, equity, and workable implementation. - These perspectives often coexist within the same institution. - Differences between administrators, writing center leaders, and faculty can create productive debate or direct conflict. - Recognizing these mindsets helps institutions engage stakeholders more effectively. ## Institutions Need Alignment, Not More Tools - Leaders consistently asked for alignment with institutional priorities, constraints, and values—not additional technology. - Effective partners should help institutions understand trade-offs rather than impose preselected solutions. - Skeptics need language that allows them to raise concerns constructively. - Advocates for adoption must recognize that resistance often reflects responsibility rather than fear of change. - AI is forcing institutions to clarify long-standing tensions such as: - Speed versus rigor - Access versus control - Innovation versus stability ## Academic Integrity as a Trust Problem - The common starting question—how to prevent students from misusing AI—is too narrow. - Academic integrity also asks whether institutions trust students and whether students trust their institutions. - Excessive restrictions can communicate distrust, while a lack of governance can appear negligent. - K–12 and higher education face different accountability structures, but both must create guidelines that reflect their actual educational values. - Many educators are shifting: - From detection to judgment - From surveillance to discernment - From punishment to responsibility - Integrity policies therefore communicate what an institution believes learning is for. ## Governing Under Uncertainty - Leaders are tired of portraying AI as either an existential threat or a universal solution. - They need principled language for discussing uncertainty with students, faculty, families, and governing boards. - Every AI decision sends a message about institutional values and credibility. - Maintaining trust requires thoughtful governance, shared understanding, and honest engagement with uncertainty—not stricter rules alone. Institutions should treat AI governance as an ongoing process of alignment and trust-building. Rather than beginning with enforcement or technology procurement, they should clarify their values, acknowledge competing perspectives, and develop policies that support informed judgment and shared responsibility.

Read original(opens in new tab)
google3 min readCurated summary

Towards developing future-ready skills with generative AI

Vantage is a Google Research experiment that uses generative AI to assess durable “future-ready” skills such as critical thinking, collaboration, conflict resolution, and creativity. It places students in realistic conversations with AI teammates, dynamically introduces challenges, and evaluates performance against educational rubrics. A study with New York University found that AI-generated scores agreed with human expert ratings at a comparable level to agreement between human raters. ## Why Future-Ready Skills Are Difficult to Measure - Skills such as collaboration, creative thinking, and conflict resolution are increasingly important as technology changes work and education. - Traditional tests are too rigid to capture how people think, communicate, and respond in realistic situations. - Human-based assessments can be resource-intensive, difficult to standardize, and dependent on whether challenging situations arise naturally. - Vantage aims to make these skills measurable, scalable, and useful for guiding instruction and student growth. ## AI-Simulated Team Assessments - Students participate in open-ended tasks, such as preparing a debate or pitching a creative idea, alongside AI avatars. - An “Executive LLM” uses an assessment rubric to manage the conversation and introduce targeted challenges, such as disagreement or conflict. - This adaptive process is designed to elicit enough evidence to assess a particular skill while keeping the interaction natural. - An “AI Evaluator” reviews the conversation transcript using the same rubric. - Students receive a visual skill map and qualitative feedback describing their demonstrated strengths and areas for improvement. ## Validation with New York University - Google Research partnered with NYU to align Vantage’s tasks and scoring criteria with established educational rubrics. - The joint study involved 188 U.S. participants aged 18–25 and focused on conflict resolution and project management. - Researchers tested whether the Executive LLM could steer conversations toward specific skills. - Steered conversations produced significantly more skill-relevant information than conversations involving independent, uncoordinated AI avatars. - The AI Evaluator’s scores showed agreement with human expert ratings comparable to the agreement between two human raters. - The results suggest that LLM-based assessment can provide a scalable alternative for evaluating complex interpersonal skills. ## Additional Research - Google also collaborated with OpenMic to study creativity and English language arts. - The collaboration analyzed work from 180 students completing creative multimedia assignments, including character interviews and literature-related media articles. - These studies tested whether the evaluation approach could extend beyond collaboration-focused tasks. Vantage is available in English through Google Labs as a research experiment. Its approach could help educators provide more consistent practice, evidence-based feedback, and scalable assessment for skills that conventional tests struggle to capture.

Read original(opens in new tab)
gitlab3 min readCurated summary

GitLab named a 2026 Omdia Universe Leader

GitLab was named a Leader in Omdia’s 2026 Universe for AI-assisted Software Development, IDE-based Tools, ranking among 19 vendors. Its strongest results came from covering the entire software lifecycle—not just code generation—including planning, security, testing, deployment, and operations. The report suggests that AI delivers the greatest productivity gains when automation extends beyond coding into coordinated, governed delivery. ## Omdia’s Broader Evaluation - Omdia expanded its criteria to assess full software lifecycle capabilities. - The report emphasized that faster coding alone can create downstream bottlenecks in: - Code review - Security remediation - Testing - Deployment coordination - Agentic AI was evaluated as a current capability, including: - Autonomous task coordination - Handoffs between specialized agents - Support for teams at different stages of AI adoption - Omdia categorizes vendors as Leaders, Challengers, or Prospects based on capability and strategy/execution. ## GitLab’s Top Scores - **Solution Breadth: 100%** - Covers planning, requirements, development, security, deployment, and issue management in one platform. - Planner Agent and Security Analyst Agent extend AI into sprint planning, vulnerability triage, and remediation guidance. - **Strategy and Innovation: 88%** - Uses end-to-end orchestration and a privacy-first architecture that does not train on private customer data. - Supports multiple models through partnerships with Anthropic, Google, and AWS. - Provides shared context across issues, merge requests, pipelines, and security findings. - **Core Features: 82%** - Offers context-aware code generation, unit and integration testing, security testing, and review prioritization. - Automates CI/CD, GitOps, and pipeline-failure root cause analysis. - The AI Impact Dashboard tracks cycle time, deployment frequency, and productivity effects. - GitLab also received top-tier scores for Extended Features (80%) and Vendor Execution (88%). ## Developers and AI Agents - Teams are increasingly structured around engineers supervising AI agents. - Human responsibilities are shifting toward: - Defining requirements and guardrails - Supervising quality and security - Designing autonomous production pipelines - Connecting business objectives with agentic systems - Automating only code generation provides limited benefit if review, testing, and deployment remain manual. ## Enterprise Readiness - Omdia treated compliance, privacy, and deployment flexibility as baseline requirements for Leader-tier platforms. - GitLab highlights: - SOC 2 and ISO 27001 certification - No training on private customer data for agentic AI - Self-managed, cloud, on-premises, and air-gapped deployment - Support for self-hosted AI models - GitLab Dedicated, including FedRAMP Moderate authorization for government - These capabilities target regulated industries requiring strong data residency, auditability, and governance. GitLab’s central argument is that AI coding speed matters only when the rest of the software delivery lifecycle can keep pace. Engineering teams should evaluate AI platforms by their ability to deliver secure, governed, production-ready software—not merely by how much code they can generate.

Read original(opens in new tab)