AI Agents

171 posts

aws2 min readCurated summary

Modernize your workflows: Amazon WorkSpaces now gives AI agents their own desktop (preview) | Amazon Web Services

Amazon WorkSpaces now lets AI agents operate desktop and legacy applications directly, eliminating the need to build APIs or modernize existing software. Agents use managed virtual desktops with IAM authentication, security controls, and auditability through CloudTrail and CloudWatch. The feature is in public preview and supports agent frameworks through the Model Context Protocol (MCP). ## The Challenge of Legacy Applications - Many enterprises depend on applications without modern APIs: - 75% of organizations reportedly run legacy applications. - 71% of Fortune 500 companies rely on mainframe-based processes with limited programmatic access. - Organizations can either delay AI adoption or undertake costly, risky modernization projects. ## AI Agents in Secure WorkSpaces - AI agents operate desktop applications inside managed WorkSpaces environments. - Agents authenticate with AWS Identity and Access Management (IAM). - Existing security and compliance controls remain in place because agents do not run on local machines. - AWS CloudTrail and Amazon CloudWatch provide audit trails. - WorkSpaces supports MCP, making it compatible with frameworks such as LangChain, CrewAI, and Strands Agents. ## Configuring Agent Access - Administrators create a WorkSpaces Applications stack and enable the **Add AI Agents** option. - Agent capabilities can include: - **Computer input:** Clicking, typing, and scrolling. - **Computer vision:** Capturing screenshots so the agent can interpret the interface. - **Screenshot storage:** Saving session images for auditing and debugging. - Administrators define screen resolution and image format. The example uses 1280×720 resolution and PNG images. - Agents connect through a managed MCP endpoint using IAM credentials. ## Automating Unmodified Desktop Workflows - A Strands Agent SDK and Amazon Bedrock example completes a prescription refill by: - Looking up a patient record. - Searching for medication. - Placing the order. - Confirming the refill. - The pharmacy application requires no API, code changes, migration, or awareness that an agent is controlling it. ## Availability - The feature is in public preview at no additional cost. - It is available in selected AWS Regions across the United States, Canada, Europe, and Asia. - Developers can begin with AWS’s GitHub repository or the Amazon WorkSpaces product page. Organizations can use WorkSpaces as a governed execution environment for AI agents, allowing them to automate legacy desktop workflows while postponing or avoiding extensive application modernization.

Read original(opens in new tab)
gitlab3 min readCurated summary

8 Agentic AI patterns reshaping team collaboration

A synthesis of 17 agentic AI platforms identifies eight patterns that help teams work faster, work smarter, and maintain control. The strongest platforms do more than provide capable individual agents: they reduce coordination overhead, embed agents in existing workflows, and provide governance across the software lifecycle. The post argues that integrated team collaboration and managed deployment will distinguish leading AI platforms. ## Eight Collaboration Patterns ### Proactive Status Updates - Agents generate progress summaries from live task data. - They identify blockers, risks, and slipping deadlines before escalation. - Automated updates reduce status meetings and manual check-ins. ### Intelligent Work Routing - Agents assign work based on skills, capacity, and project context. - Continuous workload balancing replaces periodic planning adjustments. - Transparent routing logic lets humans review and correct assignments. ### Team Communication Support - Agents summarize chats, threads, and meetings. - Decisions and conversation history remain available to new participants. - Async summaries reduce repeated explanations and unnecessary meetings. ### Role-Specific Agents in Chat - Specialist agents operate inside tools such as Slack. - They can handle onboarding, IT, sales, and other role-based tasks. - Simple interactions, such as an emoji reaction, can create tracked work. ### Shared Conversational Context - Agents retain awareness of participants, threads, and files. - Teams benefit from knowledge gathered through another member’s prompt. - Shared context prevents duplicated prompting and helps new members continue work immediately. ### Role-Based Access Control - Agents inherit permissions from their assigned identities and roles. - Access controls can apply at the field level, preventing unauthorized reading or actions. - Detailed action logs provide an auditable record for compliance. ### Governed Environments - Agents move through development, testing, and production using managed pipelines. - Sandboxes isolate early development and prevent conflicts. - Controlled promotion prevents untested agents or disruptive updates from reaching production. ### Collaborative Agent Development - Multiple team members can co-own, edit, debug, and maintain agents. - Tiered permissions support shared ownership without removing accountability. - Standardized protocols help agents created by different contributors work together. ## Lessons from the Competitive Landscape - Agents are increasingly embedded in existing communication and work tools rather than isolated portals. - Governance becomes essential as organizations scale agent usage. - Agent development is evolving into a collaborative discipline requiring shared ownership, versioning, and auditing. - The biggest opportunity is reducing the “coordination tax” of meetings, check-ins, and repeated explanations. - Few platforms provide an end-to-end governance experience combining environment grouping, shared catalogs, and managed promotion pipelines. ## Implications for GitLab GitLab’s integrated DevSecOps lifecycle gives it a structural advantage because software delivery workflows, context, and controls already exist in one platform. GitLab Duo Agent Platform is positioned to embed agents directly into those workflows, allowing teams to orchestrate work while agents execute across the software development lifecycle. Teams evaluating agentic AI should prioritize not only agent capability, but also shared context, transparent automation, permissions, deployment governance, and collaborative maintenance.

Read original(opens in new tab)
aws4 min readCurated summary

AWS Weekly Roundup: What’s Next with AWS 2026, Amazon Quick, OpenAI partnership, and more (May 4, 2026) | Amazon Web Services

The AWS weekly roundup highlights a major shift toward agentic AI across Amazon’s products and its partnership with OpenAI. The biggest announcements include expanded Amazon Quick capabilities, four specialized Amazon Connect solutions, and OpenAI models and Codex becoming available through Amazon Bedrock. AWS also introduced new EC2 instances, agent optimization tools, Ruby 4.0 support for Lambda, and a transition plan from Amazon Q Developer to Kiro. ## What’s Next with AWS 2026 - AWS and OpenAI executives presented new ways businesses are using AI agents to automate operations. - The announcements centered on Amazon Quick, Amazon Connect, and deeper integration with OpenAI through Amazon Bedrock. ## Amazon Quick Expands Beyond Chat - A new desktop app, currently in preview, connects Quick to local files, calendars, and communications without requiring a browser. - Users can sign up with a personal email or Google, Apple, GitHub, or Amazon credentials; an AWS account is not required. - Quick can generate: - Documents - Presentations - Infographics - Images - New integrations include Google Workspace, Zoom, Airtable, Dropbox, and Microsoft Teams. - The preview “Build custom apps with Quick” feature lets users create intelligent applications, dashboards, and web pages using natural-language instructions. ## Amazon Connect Becomes Four Agentic AI Products - **Amazon Connect Decisions** applies Amazon’s operational expertise and supply-chain tools to help organizations move from reactive crisis management to proactive planning. - **Amazon Connect Talent** provides AI-led interviews, science-backed assessments, and consistent candidate evaluations for large-scale hiring. - **Amazon Connect Customer**, the renamed customer-service product, supports personalized voice, chat, and digital experiences, with conversational AI that can be configured in weeks. - **Amazon Connect Health** supports patient verification, appointment management, patient insights, ambient documentation, and medical coding. ## OpenAI Partnership Expands Through Amazon Bedrock - OpenAI models, including GPT-5.5 and GPT-5.4, are coming to Bedrock in limited preview. - Customers can use existing Bedrock APIs with AWS security, governance, and cost controls, without managing new infrastructure. - **Codex on Amazon Bedrock** brings OpenAI’s coding agent into AWS environments: - Authentication uses AWS credentials. - Inference runs through Bedrock. - Usage can count toward AWS cloud commitments. - Initial access includes the Codex CLI, desktop app, and Visual Studio Code extension. - **Bedrock Managed Agents powered by OpenAI** combines OpenAI models with AWS infrastructure and the OpenAI harness for long-running, production-oriented agent workflows. ## New EC2 Instance Families - **M8in and M8ib** instances are generally available, offering up to 43% higher performance than M6in and M6ib. - M8in provides up to 600 Gbps of network bandwidth. - M8ib provides up to 300 Gbps of EBS bandwidth. - **R8in and R8ib** target memory-intensive workloads such as commercial databases, data lakes, and SAP HANA. - **C8ine and M8ine** provide up to 2.5 times higher packet performance per vCPU and up to twice the internet-gateway throughput of their predecessors. - These network-optimized instances are designed for virtual firewalls, load balancers, security appliances, and 5G user-plane workloads. ## AgentCore and Lambda Updates - Bedrock AgentCore Optimization, in preview, adds: - Production-trace analysis - Recommendations for system prompts and tool descriptions - Batch evaluations - A/B testing against live traffic - Recommendations require human approval before deployment. - AWS Lambda now supports Ruby 4.0 as a managed runtime and container base image. - Ruby 4.0 support includes advanced logging features such as structured JSON logs, configurable log levels, and custom CloudWatch log groups. ## Amazon Q Developer Moves Toward Kiro - Amazon Q Developer IDE plugins and paid subscriptions will reach end of support on April 30, 2027. - New signups will be blocked beginning May 15, 2026. - Existing subscriptions can continue adding users until then. - Opus 4.6 will leave Q Developer Pro on May 29, 2026, while newer coding models such as Opus 4.7 will be exclusive to Kiro. - Q Developer experiences in the AWS Console, documentation, mobile app, Slack, and Microsoft Teams are unaffected. AWS’s direction is increasingly centered on managed AI agents integrated into everyday business workflows. Organizations adopting these services should evaluate the new Bedrock, Quick, and Connect capabilities while also planning migration from Q Developer to Kiro before the announced support deadlines.

Read original(opens in new tab)
figma2 min readCurated summary

Workflow Lab: Expanding the Canvas with Figma MCP | Figma Blog

Figma’s workflow demonstrates how the Figma MCP server can reconnect design and implementation as features evolve. By reading coded states and generating editable frames on the canvas, an agent exposes product behavior that was invisible in the original design. This lets designers improve real edge cases and compare the shipped experience with design intent earlier. ## The problem: Code creates new product states - Astra, a fictional AI video platform, ships features rapidly with agentic coding tools. - An initial export flow covered sequence selection, format choice, settings confirmation, and export. - As development progressed, additional states appeared: - Encoding errors - Rendering and loading progress - Empty selections - Unsupported formats - These states were not necessarily design oversights; they emerged from real code and data. - When the canvas represents only the initial flow, designers cannot fully address the experience users will encounter. ## Expanding the canvas with Figma MCP - The Figma MCP server allows an agent to read implementation details and write results to the Figma canvas. - Using `use_figma`, the agent identifies coded states and creates editable frames using the team’s design-system components. - Astra’s canvas expands from four original frames to fourteen frames representing the broader product reality. - This replaces a slower task-and-ticket feedback loop with a direct conversation between design, code, and the agent. ## Designing better edge cases - The designer can immediately work on states that previously remained hidden: - Adds recovery guidance to the encoding error state. - Enhances the render loading state with progress information and an estimated completion time. - Adds copy and personality to the empty-selection state to encourage feature adoption. - Designers spend less time discovering missing requirements and more time shaping actual product behavior. - The canvas becomes a shared workspace for reviewing the full experience, not merely documenting the initial concept. ## Comparing design and implementation - The workflow also places the coded version beside the original Figma design for visual comparison. - A findings panel surfaces discrepancies by severity. - Example differences include: - A larger modal title - An additional “Post share link” button - A removed settings-panel surface - A demoted settings header The practical recommendation is to use Figma MCP as an ongoing design-code feedback loop: bring real implementation states onto the canvas, refine them with design expertise, and use visual comparisons to catch drift before it becomes part of the shipped product.

Read original(opens in new tab)
figma3 min readCurated summary

FigJam Is Now Your Coding Agent’s Whiteboard Too | Figma Blog

FigJam is being positioned as a shared whiteboard for coding agents and engineering teams. New MCP skills let agents generate architecture and ER diagrams, write to and read from FigJam, and turn research or project plans into collaborative visual boards. The workflow connects agent-generated planning, human review, and implementation, reducing architectural confusion as teams ship code faster. ## Turning Agent Output into Visual Plans - The author built on Figma’s existing `generate_diagram` MCP tool to support more complex architecture and ERD layouts. - The new `figma-use-figjam` MCP skill allows agents to read and write directly to FigJam boards. - Skills such as `generate-project-plan` can transform documentation, codebases, and conversations into visual project plans. - Diagrams can include: - Architecture and entity-relationship diagrams - Notes and annotations - Code blocks - Implementation context and technical decisions ## Step 1: Research, Plan, and Visualize - The coding agent gathers relevant documentation, codebase structure, existing patterns, and implementation constraints. - It evaluates possible solutions, researches tradeoffs, identifies affected services and files, and proposes stacked PRs and testing strategies. - Instead of leaving the plan in a dense Markdown document, the agent exports it to FigJam as an interactive architecture review. - Visualizing the options helps teams understand the system and identify the cleanest approach more quickly. ## Step 2: Collaborate Before Coding - Engineers share the FigJam board with teammates for asynchronous or live review. - Team members can comment on concrete design questions, such as: - Whether a tool should support multiple file types - Whether it should accept a `folderId` - Where newly created files should be stored - FigJam provides a collaborative format that preserves technical context for distributed teams. - Teams can review and refine agent-generated diagrams before implementation begins. ## Step 3: Feed Decisions Back to the Agent - After review, the author uses the `get_figjam` tool to retrieve the board’s diagrams, comments, and decisions. - The coding agent uses that context to update the implementation plan and begin coding. - Pull requests can link back to the FigJam board, preserving the architectural rationale alongside the code. - Because the design has already been reviewed, the resulting PR is easier to evaluate and merge. ## Broader Figma Integration - The workflow builds on `use_figma`, which lets agents create or edit designs directly on the Figma canvas using real components. - `create_new_file` allows agents to generate designs in new Figma files. - Together, these capabilities extend agent collaboration beyond code into design, architecture, planning, and technical communication. Teams adopting coding agents can use FigJam as a reviewable source of shared context: let agents generate the initial plan, have humans refine the architecture visually, then return the approved decisions to the agent for implementation.

Read original(opens in new tab)
kakao3 min readCurated summary

Kanana Scala 1st Seminar On-site Sketch

Kakao’s first Kanana Scholar seminar brought together seven leading AI professors and Kakao researchers to discuss the company’s independent AI strategy. Kakao presented its from-scratch Kanana foundation models, emphasizing data efficiency, Korean-language capability, and multimodal processing. The discussion concluded that Kakao should focus less on generic benchmark scores and more on technology sovereignty, personalized agents, and practical execution in real services. ## Kanana Foundation Models - Kakao is developing its own foundation-model lineup to strengthen competitiveness and reduce dependence on overseas providers. - Kanana reportedly achieved strong performance using 11 trillion training tokens, compared with 23 trillion tokens for a similarly sized global-target model. - Kakao attributed this efficiency to the quality and refinement of its training data. - The company also demonstrated **Kanana-o**, an omni model capable of processing text, images, and audio in real time. - The model handled emotional speech and multi-speaker conversations naturally, receiving praise for its Korean fluency. ## Technology Sovereignty and Customization - Kakao argued that proprietary models protect it from external risks such as changing licensing policies and closed technologies. - Owning the technology enables Kakao to build efficient, customized models optimized for its services. - Participating professors agreed that control over Korean cultural context and local issues is essential for technological sovereignty. - They viewed an independent model as a strategic asset for long-term service stability. ## Digital World Models and Personalized Agents - Kakao aims to understand users’ behavioral context within KakaoTalk and provide highly personalized assistance. - On-device AI could protect private conversations while allowing agents to respond immediately to user needs. - The professors suggested expanding the idea of “physical AI” into a **digital world model** that predicts interactions and causal relationships across a platform. - This direction could create an area of AI differentiation uniquely suited to Kakao’s ecosystem. ## Evaluating Practical Agentic Intelligence - Kakao is prioritizing AI systems that can create multi-step plans, call necessary tools, and complete tasks independently. - It plans to use an internally developed orchestration benchmark to evaluate real-world problem-solving ability. - The professors cited Claude as an example of how users perceive intelligence through successful completion of complex requests, not merely high benchmark scores. - They recommended competing through practical execution in real service environments rather than focusing only on text-generation performance. ## Industry-Academic Cooperation - Kakao plans to explore GPU support for university research labs and undergraduate AI clubs. - Possible support could include credits, project-based resources, and other forms of infrastructure assistance. - The seminar marked the beginning of a broader collaboration aimed at advancing Korea’s AI ecosystem and developing future talent. Kakao’s recommended path is to combine proprietary, efficient models with privacy-preserving personalization and strong agentic execution. Success will depend on how effectively Kanana turns technical depth into useful intelligence that users can experience in everyday services.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Moving past bots vs. humans

The distinction between bots and humans is becoming too blurry to serve as the foundation of web protection. Browsers, accessibility tools, proxies, and AI agents can all behave differently while representing legitimate users, while human activity can also be malicious. Website owners should instead focus on intent, behavior, resource usage, and trust. ## The Web’s Original Balance - Browsers act as user agents, mediating between people and websites. - Websites rely on browser conventions to: - Present content correctly across devices. - Support purchases, logins, media, and accessibility. - Deliver advertising and control user experiences. - The web has historically balanced publisher interests with user freedoms through browser standards, extensions, and accessibility requirements. - AI agents disrupt this balance by fetching raw content without rendering pages like browsers. - Publishers often cannot tell whether a request supports one private summary or large-scale model training, making traffic and monetization less predictable. ## The Client-Server Model - Clients request resources from servers, which respond with the requested content. - Websites can scale through additional servers, caching, and CDNs. - The model’s openness allows many types of clients to interact with servers without requiring servers to understand their internal software. - That flexibility creates uncertainty: servers generally cannot see whether a response is: - Rendered for one person using a browser. - Automatically collected, archived, indexed, or reused by another system. ## Why Bot Management Exists - Websites must decide which requests they can afford to serve when capacity, CPU, or cost limits are reached. - Randomly dropping requests is possible but risks blocking legitimate users. - Access controls are also used to: - Separate attacks from normal traffic. - Manage non-malicious load. - Prevent data extraction and fake account creation. - Limit ad fraud and automated actions. - Web clients are unauthenticated by default, so services infer identity and intent from partial signals such as request volume and IP addresses. - A high-volume IP may indicate abuse, a VPN, or multiple users sharing one address, making simple bot-versus-human classifications unreliable. ## Toward Intent and Behavior-Based Protection - The important questions are whether traffic represents an attack, whether crawling is proportional to returned traffic, whether a login from a new country is expected, or whether advertisements are being manipulated. - “Bots” encompass two separate concerns: - Whether known crawlers should receive access when they provide little traffic or value in return. - Whether emerging clients behave unlike traditional browsers, affecting systems such as private rate limits. - Automation detection remains necessary, but protection systems should be designed for a future where automation is common among both legitimate and malicious actors. Website protection should evolve from identifying “bots” to evaluating intent, behavior, proportionality, and risk. The goal is not to determine whether a client is human, but whether its activity is expected, sustainable, and trustworthy.

Read original(opens in new tab)
figma2 min readCurated summary

How AI Leaders Are Borrowing From the Design Playbook | Figma Blog

AI transformation requires more than deploying new tools; it requires redesigning how organizations work. Figma argues that the most effective AI leaders adopt design practices—hands-on experimentation, close observation of workflows, and rapid prototyping—to turn adoption and innovation into meaningful business change. ## AI Leadership as Organizational Design - New AI innovation and acceleration roles are emerging to improve workflows, speed product launches, and expand tool adoption. - These leaders often coordinate AI strategy across product, support, internal operations, and technology investments. - A major risk is “performative progress”: adopting tools for appearances without changing the underlying systems and processes. - Effective leaders connect technology, teams, workflows, and business outcomes. ## Learn the Material by Using It Yourself - Leaders need firsthand experience with AI tools rather than relying only on strategic or executive-level perspectives. - Prompting, building agents, and experimenting across different tools reveals practical limitations, trade-offs, and adoption barriers. - Personal projects—such as planning travel, organizing events, or managing volunteer work—can provide low-risk opportunities to develop AI fluency. - Leaders cannot effectively guide organizations through probabilistic technologies without understanding how those technologies behave in real situations. ## Observe How Teams Actually Work - Understanding AI use across the business requires studying workflows, not just tools and their outputs. - Useful signals include Slack discussions, survey responses, usage patterns, frustrations, and points where employees get stuck. - An automation may appear successful technically but fail because it adds friction to an already complicated process. - When adoption stalls, teams may be routing around the official solution and creating unofficial alternatives; observing this behavior helps identify the real problem. ## Turn Ideas Into Prototypes - Ideas often fail because teams cannot visualize or evaluate them, not because the ideas themselves are flawed. - Prototyping converts abstract AI concepts into tangible experiences that teams can discuss and test. - Tools such as Figma Make can help leaders and teams explore concepts quickly and make early possibilities easier to understand. - Design combines observation with action: leaders should learn from real behavior, then use prototypes to test potential solutions. AI leaders should therefore combine technical curiosity with design discipline: use the tools personally, study how people work, and prototype proposed changes before attempting broad implementation.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Introducing the Agent Readiness score. Check to see if your site is agent-ready

Cloudflare argues that websites must evolve beyond browser and search-engine compatibility to become usable by AI agents. Its new Agent Readiness score evaluates whether sites support standards for discovery, content access, bot control, and agent capabilities. Early data shows adoption is extremely low, creating both a challenge and an opportunity for sites that adopt these standards early. ## Agent readiness across the web - Cloudflare analyzed the 200,000 most visited domains, excluding categories unlikely to need agent interaction. - The resulting Cloudflare Radar dataset tracks adoption of AI-agent standards and will be updated weekly. - robots.txt exists on 78% of sites, but most files target traditional search crawlers rather than AI agents. - Only 4% of sites declare AI usage preferences through Content Signals. - Just 3.9% support Markdown content negotiation via `Accept: text/markdown`. - MCP Server Cards and API Catalogs based on RFC 9727 appear on fewer than 15 sites, showing how early these standards remain. ## The Agent Readiness score Site owners can test their websites at **isitagentready.com**. Cloudflare scans the site and scores it across four dimensions: - **Discoverability:** robots.txt, sitemap.xml, and Link Headers under RFC 8288. - **Content:** Markdown for Agents. - **Bot Access Control:** Content Signals, AI-specific robots.txt rules, and Web Bot Auth. - **Capabilities:** Agent Skills, API Catalogs, OAuth discovery standards, MCP Server Cards, and WebMCP. - The tool also checks commerce standards such as x402, Universal Commerce Protocol, and Agentic Commerce Protocol, though these do not yet affect the score. - Each failed check includes a prompt that can be handed to a coding agent for implementation. The service itself supports agents through a stateless MCP server with a `scan_site` tool and publishes Agent Skills documents explaining how to implement each supported standard. ## Discoverability for AI agents - robots.txt helps agents understand crawl permissions and locate sitemaps. - Sitemaps provide a structured list of site paths, reducing the need to discover content by following every HTML link. - HTTP Link headers, defined by RFC 8288, expose important resources directly in responses without requiring agents to parse page markup. - Sites can use headers such as `rel="api-catalog"` to point agents toward machine-readable capabilities. ## Making content easier to read - `llms.txt` provides an LLM-oriented reading list at the site root, describing the site and linking to important content in a format designed for model context windows. - Markdown content negotiation lets agents request a clean Markdown version of a page with `Accept: text/markdown`. - Cloudflare measured token reductions of up to 80% compared with HTML, improving speed, cost, and the likelihood that agents can consume an entire document within their context limits. Cloudflare’s recommendation is to evaluate sites with the Agent Readiness tool and adopt the relevant standards incrementally. With current adoption so low, early support can make a site significantly easier for AI agents to discover, understand, authenticate with, and use.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Browser Run: give your agents a browser

Cloudflare is renaming Browser Rendering to Browser Run and positioning it as a full browser platform for AI agents. It provides remotely hosted Chrome sessions that agents can control, observe, debug, record, and scale globally, while allowing humans to intervene when necessary. The update expands access through CDP and MCP, making existing automation tools and AI coding assistants compatible with Cloudflare’s browser infrastructure. ## Browser Run for AI Agents - Agents can navigate websites, read content, fill out forms, extract data, take screenshots, and verify results. - Browser sessions run on Cloudflare’s global network, reducing infrastructure and browser-maintenance requirements. - Sessions can scale dynamically and open near users for lower latency. - The platform now supports up to 120 concurrent browsers, up from 30. ## Observability and Human Intervention - **Live View** shows an agent’s browser session in real time, making it easier to confirm success or diagnose failures. - **Human in the Loop** allows agents to transfer control when they encounter login screens or unusual edge cases. - A human can resolve the issue and return control to the agent. - **Session Recordings** capture DOM changes, interactions, and navigation for debugging and postmortem analysis. ## Browser Control Options Browser Run supports several levels of automation: - Low-level control through the Chrome DevTools Protocol (CDP). - Higher-level automation with Puppeteer and Playwright. - Quick Actions for simpler tasks. - WebMCP for websites that expose agent-discoverable actions. ### Chrome DevTools Protocol - Browser Run now exposes CDP directly through a WebSocket endpoint. - Existing CDP-based frameworks, scripts, and agent tools can connect with minimal changes. - CDP provides capabilities beyond Puppeteer and Playwright, including JavaScript debugging. - Raw protocol messages can be sent directly to models, potentially reducing token usage. - Developers can connect from any language or environment without creating a Cloudflare Worker. - Self-hosted Chrome scripts can be migrated by changing the browser WebSocket URL and adding Cloudflare authentication headers. ### MCP Client Support - MCP clients such as Claude Desktop, Cursor, Codex, and OpenCode can use Browser Run as a remote browser. - Cloudflare supports the `chrome-devtools-mcp` package, which provides browser automation, debugging, and performance-analysis capabilities. - Configuration requires pointing the MCP server to Browser Run’s CDP endpoint and supplying an API token. ### WebMCP - WebMCP is intended to make websites more reliable for AI agents. - Websites can declare actions that agents can discover and call directly. - This addresses the limitations of a web originally designed primarily for human navigation. ## Overall Direction Cloudflare’s update combines hosted browser infrastructure, multiple automation interfaces, real-time visibility, replayable sessions, and human fallback. The goal is to make browser-based agents more dependable in production while avoiding the operational burden of managing Chrome infrastructure themselves. For teams building web-using agents, Browser Run offers a practical path from self-hosted or local browser automation to scalable, observable remote sessions, especially when existing CDP, Puppeteer, Playwright, or MCP tooling is already in use.

Read original(opens in new tab)
figma2 min readCurated summary

The TL;DR on MCP: Why Context Matters and How to Put It to Work | Figma Blog

MCP (Model Context Protocol) connects AI tools to the design decisions and data stored in tools like Figma. Figma argues that giving coding agents structured access to components, tokens, and layout rules produces code that better matches the intended design and design system. It also creates a two-way workflow in which developers and designers can move between code and canvas without losing context. ## MCP Connects Design and Development - Product work is increasingly iterative rather than a linear design-to-development handoff. - MCP lets AI coding tools access Figma files as structured design sources, not merely as screenshots. - Figma’s MCP server helps bring design context into code, while code-to-canvas tools can bring working interfaces back into Figma. - This keeps the broader product team involved as designs and implementations evolve. ## Why Context Matters for AI-Generated Code - Without context, an AI tool may: - Choose a color that resembles the brand color but is not linked to the correct design token. - Recreate a card instead of reusing an established component. - Flatten a complex, nested form into a single basic element. - These seemingly minor deviations accumulate across screens and components. - MCP exposes the underlying components, tokens, and layout decisions that explain how a design was built. ## Designers: Files Directly Influence Production Code - Design systems now influence not only human implementation but also AI-generated code, from prototypes through production. - Well-structured, consistent Figma files can guide AI toward more reliable and on-brand results. - Poor organization or small inconsistencies can spread widely because AI reproduces them at scale. - MCP also lets designers review code-built interfaces in Figma, add missing states, refine details, and prepare work for production without starting over. ## Developers: Less Translation, More Building - AI coding tools can accelerate implementation, but their output is less accurate when design intent is unavailable. - MCP reduces the translation required between a visual design and working code by supplying the system and component context behind the design. - Developers can spend more time building instead of reconstructing design decisions from screenshots or incomplete handoffs. Figma’s practical recommendation is to treat design files and design systems as active inputs to AI workflows. The better the structure and context captured in those files, the more consistently AI can generate code that reflects the intended product.

Read original(opens in new tab)
github3 min readCurated summary

Hack the AI agent: Build agentic AI security skills with the GitHub Secure Code Game

Agentic AI tools can automate powerful tasks, but their autonomy creates new security risks, including prompt injection, tool misuse, memory poisoning, and compromised multi-agent workflows. GitHub’s Season 4 Secure Code Game teaches developers to recognize these threats by attacking and hardening ProdBot, a deliberately vulnerable terminal-based AI assistant. Its five levels progressively add capabilities—and corresponding attack surfaces—mirroring how real-world AI systems evolve. ## The Secure Code Game’s Evolution - The free, open-source, in-editor course teaches security by having players exploit and fix intentionally vulnerable code. - Earlier seasons covered: - General secure coding across JavaScript, Python, Go, and GitHub Actions. - LLM security, including malicious prompts and defensive techniques. - More than 10,000 developers from industry, academia, and open source have participated. - Season 4 shifts focus from AI that generates content to AI that independently browses, uses tools, calls APIs, and acts for users. ## Why Agentic AI Security Is Urgent - Agentic systems are moving rapidly from research projects into production environments. - The OWASP Top 10 for Agentic Applications identifies threats such as: - Goal hijacking - Tool misuse - Identity abuse - Memory poisoning - A Dark Reading poll found that 48% of cybersecurity professionals expect agentic AI to be the leading attack vector by the end of 2026. - Cisco reported that although 83% of organizations planned to deploy agentic AI, only 29% felt prepared to secure it. - The article argues that learning to think like an attacker is essential for closing this readiness gap. ## ProdBot: A Deliberately Vulnerable AI Assistant - ProdBot is a terminal-based productivity and coding assistant inspired by tools such as OpenClaw and GitHub Copilot CLI. - It can: - Convert natural-language requests into bash commands. - Browse a simulated web. - Connect to MCP servers. - Run organization-approved skills. - Store persistent memory. - Coordinate multiple agents. - Players’ objective is to use natural language to make ProdBot reveal the contents of `password.txt`. - No prior AI or coding experience is required; all interaction takes place through the CLI. ## Five Progressive Attack Surfaces - **Level 1: Shell execution** - ProdBot runs generated bash commands in a sandbox. - The challenge is to determine whether the sandbox can be escaped. - **Level 2: Web browsing** - ProdBot reads simulated news, finance, sports, and shopping sites. - Untrusted web content introduces risks such as instruction hijacking and prompt injection. - **Level 3: MCP integrations** - ProdBot gains access to external tool providers for stock quotes, browsing, and cloud backup. - Additional tools increase both functionality and opportunities for abuse. - **Level 4: Skills and memory** - Organization-approved plugins and persistent memory create layered trust relationships. - The level tests whether trusted skills and stored information are actually safe. - **Level 5: Multi-agent orchestration** - ProdBot combines six specialized agents, three MCP servers, three skills, and a simulated open-source project. - Claims that agents are sandboxed and data is pre-verified become assumptions to test rather than guarantees. ## Real-World Relevance - The game’s vulnerabilities reflect active security concerns in deployed autonomous AI systems rather than purely theoretical exercises. - The article cites CVE-2026-25253, known as “ClawBleed,” an OpenClaw vulnerability rated CVSS 8.8. - The flaw allowed attackers to steal authentication tokens through a malicious link and gain full control of an OpenClaw instance. - Season 4’s broader goal is to develop instincts for identifying similar weaknesses during architecture reviews, tool-integration audits, and production deployments. Developers working with AI agents should treat every new capability—shell access, browsing, plugins, memory, or collaboration—as a potential attack surface. Practicing these failure modes in a controlled environment like the Secure Code Game can help teams design safer agentic systems before deploying them.

Read original(opens in new tab)
slack4 min readCurated summary

Managing context in long-run agentic applications

Long-running multi-agent applications cannot rely on unlimited conversation history: model APIs are stateless, and growing context windows eventually reduce quality or hit hard limits. Slack’s security-investigation system addresses this by giving agents complementary, purpose-specific context rather than exposing every agent to the full investigation history. Its three main channels—the Director’s Journal, Critic’s Review, and Critic’s Timeline—preserve coherence while leaving room for independent reasoning. ## The Challenge of Long-Run Coherence - Agent frameworks usually maintain continuity by resending the complete message history with every inference request. - Long investigations can involve hundreds of requests and megabytes of generated output. - Context windows impose both: - A hard limit on how much history can be supplied. - A quality limit, because performance may degrade before the window is completely full. - Multi-agent systems need carefully scoped views: - Too little shared context makes agents disconnected from the investigation. - Too much shared context can suppress creativity and encourage confirmation bias. ## Three Complementary Context Channels Slack uses separate information sources for different purposes: - **Director’s Journal** - Structured working memory for the orchestrating Director. - Records decisions, observations, findings, questions, actions, and hypotheses. - **Critic’s Review** - An annotated report evaluating Expert findings. - Includes credibility scores to distinguish reliable evidence from weaker claims. - **Critic’s Timeline** - A consolidated chronological view of findings. - Also attaches credibility scores, helping agents understand the sequence and evidential strength of events. Together, these channels provide continuity without forcing every agent to process the entire raw conversation. ## The Director’s Journal The Director coordinates the investigation by choosing questions, assigning specialist Experts, assessing progress, and deciding when to stop. The Journal gives it persistent working memory across phases and rounds. - The Director is encouraged to update the Journal frequently with short notes. - Entries can represent: - **Decisions** about investigative strategy - **Observations** about emerging patterns - **Findings** representing confirmed facts - **Questions** that remain unresolved - **Actions** taken or planned - **Hypotheses** about what may be happening - Entries can also include: - Priority levels - Follow-up actions - References to supporting evidence - Investigation phase, round number, and timestamp - The journaling tool itself simply accumulates entries; the agents’ prompts explain how to interpret them. ## Maintaining Alignment Across Agents - The Journal creates a shared narrative around the Director’s evolving plan. - It helps the Director: - Track progress - Identify dead ends - Revise investigative direction - Preserve decisions between rounds - Guide other agents toward a conclusion - Every agent receives the current Journal chronologically, along with instructions describing: - The Director’s role - Each agent’s relationship to the Director - The Journal’s purpose - How its entries should influence their work - This approach keeps specialists anchored to the overall investigation without requiring them to read every prior interaction. ## Example Investigation Context The sample Journal comes from an investigation into an apparent kernel-module-loading alert that turned out to be a false positive. - The Director recorded that: - The event originated from a package-installation hook rather than a direct `modprobe` command. - The host appeared to be a personal development workstation. - Root access was expected in that environment. - The detection rule matched “kmod” in a script path rather than confirming module loading. - The Director identified relevant Expert domains, including: - Endpoint telemetry - Identity and access - Configuration management - User behavior - The Journal captured both the preliminary conclusion and remaining verification tasks, such as checking the parent process chain. The design therefore preserves the reasoning trail while keeping it structured and compact. A practical design for long-running agentic systems is to replace indiscriminate transcript accumulation with multiple, curated context channels. Persistent journals can maintain leadership and continuity, while independent reviews and timelines provide evidence-focused context without overwhelming agents or biasing their reasoning.

Read original(opens in new tab)
gitlab3 min readCurated summary

GitLab named a 2026 Omdia Universe Leader

GitLab was named a Leader in Omdia’s 2026 Universe for AI-assisted Software Development, IDE-based Tools, ranking among 19 vendors. Its strongest results came from covering the entire software lifecycle—not just code generation—including planning, security, testing, deployment, and operations. The report suggests that AI delivers the greatest productivity gains when automation extends beyond coding into coordinated, governed delivery. ## Omdia’s Broader Evaluation - Omdia expanded its criteria to assess full software lifecycle capabilities. - The report emphasized that faster coding alone can create downstream bottlenecks in: - Code review - Security remediation - Testing - Deployment coordination - Agentic AI was evaluated as a current capability, including: - Autonomous task coordination - Handoffs between specialized agents - Support for teams at different stages of AI adoption - Omdia categorizes vendors as Leaders, Challengers, or Prospects based on capability and strategy/execution. ## GitLab’s Top Scores - **Solution Breadth: 100%** - Covers planning, requirements, development, security, deployment, and issue management in one platform. - Planner Agent and Security Analyst Agent extend AI into sprint planning, vulnerability triage, and remediation guidance. - **Strategy and Innovation: 88%** - Uses end-to-end orchestration and a privacy-first architecture that does not train on private customer data. - Supports multiple models through partnerships with Anthropic, Google, and AWS. - Provides shared context across issues, merge requests, pipelines, and security findings. - **Core Features: 82%** - Offers context-aware code generation, unit and integration testing, security testing, and review prioritization. - Automates CI/CD, GitOps, and pipeline-failure root cause analysis. - The AI Impact Dashboard tracks cycle time, deployment frequency, and productivity effects. - GitLab also received top-tier scores for Extended Features (80%) and Vendor Execution (88%). ## Developers and AI Agents - Teams are increasingly structured around engineers supervising AI agents. - Human responsibilities are shifting toward: - Defining requirements and guardrails - Supervising quality and security - Designing autonomous production pipelines - Connecting business objectives with agentic systems - Automating only code generation provides limited benefit if review, testing, and deployment remain manual. ## Enterprise Readiness - Omdia treated compliance, privacy, and deployment flexibility as baseline requirements for Leader-tier platforms. - GitLab highlights: - SOC 2 and ISO 27001 certification - No training on private customer data for agentic AI - Self-managed, cloud, on-premises, and air-gapped deployment - Support for self-hosted AI models - GitLab Dedicated, including FedRAMP Moderate authorization for government - These capabilities target regulated industries requiring strong data residency, auditability, and governance. GitLab’s central argument is that AI coding speed matters only when the rest of the software delivery lifecycle can keep pace. Engineering teams should evaluate AI platforms by their ability to deliver secure, governed, production-ready software—not merely by how much code they can generate.

Read original(opens in new tab)
github3 min readCurated summary

GitHub Copilot CLI combines model families for a second opinion

GitHub Copilot CLI’s experimental Rubber Duck feature adds an independent reviewer from a different AI model family to catch mistakes before they compound. When Claude models orchestrate a task, GPT-5.4 reviews plans, implementations, and tests at key checkpoints. On SWE-Bench Pro, Claude Sonnet 4.6 with Rubber Duck closed 74.7% of the performance gap with Claude Opus 4.6 alone, particularly on complex, multi-file tasks. ## The Problem with Self-Review - Coding agents typically assess a task, plan, implement, test, and iterate. - Early assumptions can create downstream dependencies and make small mistakes expensive to fix. - Self-reflection helps, but a model reviewing its own work may retain the same training biases and blind spots. ## Cross-Family Review with Rubber Duck - Rubber Duck is a focused review agent powered by a complementary model family. - Claude orchestrators currently use GPT-5.4 as the reviewer. - It produces a short list of high-value concerns, including: - Missed details - Questionable assumptions - Architectural risks - Relevant edge cases ## Evaluation Results - On SWE-Bench Pro, Sonnet 4.6 plus Rubber Duck approached the resolution rate of Opus 4.6 running alone. - Benefits were strongest for problems involving at least three files and 70 or more steps. - Sonnet plus Rubber Duck scored: - 3.8% above the Sonnet baseline on difficult tasks - 4.8% higher on the hardest tasks across three trials - Examples included detecting: - A scheduler that would start and immediately exit - A loop overwriting one dictionary key and dropping Solr facet categories - Cross-file Redis references that would silently break email confirmation flows ## When Reviews Happen Rubber Duck can be invoked automatically, reactively, or on request: - After a plan is drafted, to prevent flawed decisions from spreading. - After complex implementation work, to identify edge cases. - After tests are written but before they run, to expose coverage gaps or weak assertions. - When the primary agent is stuck or repeating an unproductive loop. - Any time the user asks Copilot to critique its work. Copilot incorporates the feedback and explains what changed. Reviews are intentionally infrequent and targeted at checkpoints where they provide the most value. ## Availability and Use Cases - Rubber Duck is available in Copilot CLI’s experimental mode through `/experimental`. - It works with Claude Opus, Sonnet, and Haiku as orchestrator models, provided the user has GPT-5.4 access. - It is especially suited to: - Complex refactors and architectural changes - High-stakes coding tasks - Test coverage review - Getting a second opinion before committing to a plan Rubber Duck is a practical way to reduce model-specific blind spots by combining different AI families. Developers can enable it experimentally in Copilot CLI and use automatic or on-demand critiques for difficult work.

Read original(opens in new tab)