Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

Welcome to Agents Week

Cloudflare argues that AI agents require a fundamental shift in Internet and cloud infrastructure. Unlike traditional one-to-many applications, agents create unique, ephemeral execution environments for individual users and tasks, making current container-based economics and scaling inadequate. The company positions lightweight V8 isolates, alongside containers and browser support, as the foundation for making agents practical at global scale. ## The Internet Was Built for Applications, Not Agents - Cloud infrastructure evolved during the smartphone era to serve many users through a finite number of application instances. - Microservices, containers, Kubernetes, load balancing, and replication all support this one-to-many model. - Agents differ because an LLM dynamically determines code paths, tool usage, and task duration. ## One User, One Agent, One Task - Each agent may need its own execution environment, filesystem, tools, and state. - Coding agents currently use containers with access to Git, Bash, filesystems, and arbitrary binaries. - As agents spread to assistants, analysts, customer service, and planning tasks, the number of simultaneous environments could grow dramatically. ## The Scale Challenge - If 100 million US knowledge workers used agents at 15% concurrency, infrastructure would need about 24 million simultaneous sessions. - At 25–50 users per CPU, that implies roughly 500,000 to 1 million server CPUs in the US alone. - Multiple agents per person and global adoption would increase demand by orders of magnitude. ## Isolates as Agent Infrastructure - Cloudflare’s Workers platform uses V8 isolates instead of containers. - Isolates start in milliseconds, use only a few megabytes of memory, and provide secure sandboxing. - They can be up to 100 times faster to start and up to 100 times more memory-efficient than containers. - Dynamic Workers can create execution environments on demand, run code, and discard them at a scale of millions per second. - This efficiency could make one-agent-per-user economics viable beyond expensive coding assistants. ## The “Horseless Carriage” Phase - Early agent infrastructure often adapts existing systems instead of using designs built specifically for agents. - Agents use headless browsers to navigate human-oriented websites, though structured protocols such as MCP could provide direct service access. - Many MCP servers simply wrap REST APIs, despite LLMs often being better at writing and executing code than making long sequences of tool calls. - CAPTCHAs and behavioral fingerprinting ask whether a requester is human, while agent systems need identity, authorization, and permission controls. - Full containers are frequently used for tasks that require only a few API calls and a response. ## Supporting Both Old and New Models - Infrastructure transitions rarely happen all at once; technologies such as IPv4/IPv6, HTTP/2/HTTP/3, and TLS 1.2/1.3 coexist. - Cloudflare plans to support existing agent workloads while developing more efficient primitives. - Containers remain important for coding agents that need filesystems, Git, Bash, and arbitrary binaries. - Cloudflare is also expanding container-based sandbox environments and browser-rendering capabilities for services that do not yet support agent-native protocols. Cloudflare’s broader recommendation is to build infrastructure that can serve today’s container-based agents while moving toward lightweight, ephemeral isolates designed for billions of specialized agent sessions.

Read original(opens in new tab)
cloudflare4 min readCurated summary

500 Tbps of capacity: 16 years of scaling our global network

Cloudflare’s network has grown from a single transit provider in 2010 to more than 500 Tbps of provisioned external capacity across 330+ cities. The company argues that this scale is not merely about bandwidth: it enables security decisions, application execution, and routing validation to happen locally on every server. Its distributed architecture can absorb massive attacks automatically while supporting edge computing and emerging Internet protocols. ## From Transit Provider to Global Network - Cloudflare began with nLayer Communications as its first transit provider. - Expansion required city-by-city work: colocation contracts, fiber installation, hardware deployment, and Internet exchange peering. - In 2018, Cloudflare opened 31 cities in 24 days, despite logistical challenges such as customs delays and missing equipment. - The network now spans more than 330 cities and protects over 20% of the web. - The 500 Tbps figure represents provisioned interconnection capacity across transit, private peering, Internet exchanges, and Cloudflare Network Interconnect ports—not peak traffic. ## Turning the Network into a Security Layer - Cloudflare expanded from caching websites to securing employees and enterprise networks. - Its systems establish secure tunnels to private subnets and advertise customer IP space through BGP. - In 2025, Cloudflare mitigated a 31.4 Tbps DDoS attack lasting 35 seconds. - The attack was part of more than 5,000 attacks blocked that day, without paging an engineer. - Distributed automation allows attacks that once required nation-state resources to be handled in seconds. ## Packet-Level DDoS Mitigation - Incoming packets enter an XDP program chain in driver mode immediately after reaching the network interface card. - The `l4drop` eBPF program applies mitigation rules generated by `dosd`, Cloudflare’s denial-of-service daemon. - Each server identifies heavy traffic sources and shares the information across its colocation facility. - Mitigation rules spread globally through Quicksilver, Cloudflare’s distributed key-value store. - Only legitimate traffic reaches Unimog, the Layer 4 load balancer; Magic Transit traffic receives additional stateful inspection through `flowtrackd`. - The 31.4 Tbps attack was stopped at line rate without centralized scrubbing or human intervention. - Sufficient physical port capacity remains essential: software defenses cannot work if the network cannot first absorb the traffic. ## A Developer Platform at the Edge - Because Cloudflare already runs software on every server for packet filtering, it extended the same infrastructure to customer code through Workers. - Workers, KV, and Durable Objects run across Cloudflare’s global footprint rather than in a small number of cloud regions. - Workers Containers, introduced in 2025, support heavier workloads at the edge. - V8 isolates and custom filesystem layers reduce cold-start times. - Applications run on the same servers that discard malicious traffic before it reaches the network stack. ## Securing Routing with RPKI and ASPA - Cloudflare uses IPv6 and RPKI to reduce the risk of BGP hijacks. - It signs Route Origin Authorizations and rejects routes that fail Route Origin Validation, even when misconfigured networks become temporarily unreachable. - ASPA will extend protection by validating the network path, not just the organization authorized to originate a prefix. - The post compares RPKI to checking a destination passport and ASPA to verifying the entire flight manifest. - Cloudflare says 867,000 prefixes now have valid RPKI certificates, compared with nearly none a decade ago. - The company promotes early adoption of routing security standards because delays leave the Internet exposed to hijacks and route leaks. ## AI Agents and Internet Traffic - AI crawlers, training systems, and autonomous agents now generate more than 4% of HTML requests on Cloudflare’s network. - Human-initiated “user action” crawling increased more than 15-fold in 2025. - Unlike browsers, crawlers may retrieve every linked resource at maximum speed, making legitimate activity difficult to distinguish from attacks. - Cloudflare uses verified bot IP ranges, TLS fingerprints, behavioral analysis, and robots.txt signals to classify AI crawlers. - These signals help site owners decide which automated agents to permit. Cloudflare’s central lesson is that a global network must combine abundant capacity with intelligence distributed across every server. Its continued investment in automated mitigation, edge execution, routing security, and traffic classification is intended to make the Internet faster, safer, and more resilient as traffic patterns evolve.

Read original(opens in new tab)
netflix3 min readCurated summary

Evaluating Netflix Show Synopses with LLM-as-a-Judge

Netflix developed an LLM-as-a-Judge system to evaluate show synopses at the scale of its extensive catalog. The system assesses creative quality against expert-defined standards while also examining whether scores predict member behavior. With calibrated prompts, extended reasoning, and consensus scoring, the approach achieves more than 85% agreement with creative writers and can identify potentially impactful synopsis problems before a title launches. ## Defining a Good Synopsis - Synopsis quality is measured in two ways: - **Creative quality:** how well a synopsis follows Netflix’s editorial standards. - **Member feedback:** how the synopsis affects viewing decisions and early engagement. - Strong synopses help members quickly understand and choose titles. - Weak or misleading synopses can cause frustration, abandonment, and reduced viewing. ## Building Expert-Labeled Evaluation Data - Creative experts initially labeled roughly 1,000 diverse synopses. - Three writers scored each synopsis and explained their decisions. - Because the task was subjective, Netflix used eight calibration rounds to improve consistency. - Techniques that increased agreement included: - Replacing 1–4 ratings with binary scores. - Allowing writers to consult previous examples. - Maintaining a searchable taxonomy of recurring errors. - A model-in-the-loop process helped resolve disagreements: - Multiple writers supplied scores. - An LLM aggregated the judgments. - Writers reviewed cases with significant disagreement. - The resulting “golden set” contains about 600 synopses with criterion-level labels and explanations. ## Measuring Member Impact - Netflix uses two behavioral metrics: - **Take fraction:** how often members who see a synopsis start watching the title. - **Abandonment rate:** how often viewers stop shortly after beginning. - These metrics act as short-term proxies for long-term retention and have been validated through A/B testing. - Netflix evaluates whether LLM-generated quality scores can predict these engagement outcomes. ## Criterion-Specific LLM Judges - Initial prompts provide: - Relevant show metadata. - A summary of the applicable quality guidelines. - A request for an explanation followed by a binary score. - A single prompt covering every criterion performed poorly because it overloaded the model. - Separate judges for individual criteria performed better. - Binary outputs make evaluation straightforward using accuracy against the expert-labeled golden set. ## Improving Prompts and Reasoning - Netflix applies Automatic Prompt Optimization to a development set of about 300 examples. - Prompts are then manually refined with LLM assistance. - Performance varies significantly by criterion: prompts work well for areas such as precision but less well for subjective criteria such as clarity. - Inference-time scaling improves difficult judgments through: - **Longer rationales**, which give the model more room to reason. - **Consensus scoring**, which samples multiple judgments and combines their results. ## Tiered Rationales - Longer explanations generally improve accuracy, but they become harder for creative experts to read and audit. - Netflix therefore uses tiered rationales: - The model may reason at length internally. - It produces a concise explanation before the final score. - This approach preserves the benefits of extended reasoning while improving interpretability. - For example, the tone evaluator’s accuracy increased from 86.55% to 87.85% with tiered rationales. Netflix’s approach combines expert standards, calibrated evaluation data, specialized prompts, and inference-time reasoning to scale synopsis-quality review. The practical recommendation is to use LLM judges as carefully aligned evaluators—not generic critics—while validating their scores against both human judgment and real member behavior.

Read original(opens in new tab)
github3 min readCurated summary

GitHub Copilot CLI for Beginners: Getting started with GitHub Copilot CLI

GitHub Copilot CLI brings Copilot’s agentic coding capabilities directly into the terminal, allowing developers to inspect projects, generate code, run tests, and correct errors without switching tools. The post introduces the tool, explains installation and authentication, and demonstrates how to use it for project overviews, coding tasks, and delegated work. Its central message is that Copilot CLI can preserve development flow while supporting increasingly autonomous coding workflows. ## What GitHub Copilot CLI Does - Runs Copilot from a command-line interface with context from the current repository. - Can autonomously: - Build or modify code - Run tests - Detect and correct errors - Explore project files and documentation - Lets developers review results and request follow-up changes directly in the terminal. - Can delegate well-defined tasks to the Copilot cloud agent. ## Installing Copilot CLI - The primary cross-platform installation method, assuming Node.js is available, is: ```bash npm install -g @github/copilot ``` - Users can also install it through package managers such as Homebrew or WinGet. ## First-Time Setup - Launch the tool by entering `Copilot` in the terminal. - Authenticate with GitHub using: ```plaintext /login ``` - Authentication connects the CLI to the user’s Copilot account and the read-only GitHub MCP server. - Copilot must be granted permission to access the current folder so it can inspect or modify files. - Folder permissions can apply only to the current session or be saved for future sessions. ## Common Development Tasks - **Understand an existing project** - Prompt Copilot with: ```plaintext Give me an overview of this project ``` - It examines important files and summarizes the project structure and purpose. - **Generate new code** - For example: ```plaintext Let’s add a new endpoint to return all categories ``` - Copilot reviews existing conventions, documentation, and examples before proposing or creating files. - It requests permission before making changes. - **Delegate work to the cloud agent** - A task can be sent using: ```plaintext /delegate Let’s deal with issue #14 to add the rest of the CRUD endpoints to games ``` - The cloud agent retains the current context, creates a branch, opens a draft pull request, and performs the work in the background for later review. ## What Comes Next The broader beginner series will cover interactive mode, non-interactive mode using the `-p` flag, slash commands, and MCP server integration. These features expand Copilot CLI from an interactive coding assistant into a flexible terminal-based automation tool. Copilot CLI is recommended for developers who want AI assistance without leaving the shell: install it with npm, authenticate, grant project permissions, and begin with exploratory prompts before assigning code changes or delegated tasks.

Read original(opens in new tab)
line4 min readCurated summary

Clearing Review Bottlenecks with AI - Transforming Review Culture with PR Review Support and Internal Workshops

Orchestration Guild member Fukuyama describes how Yahoo! Places addressed PR review bottlenecks by combining AI assistance with standardized processes and team culture. Reviews had become concentrated among a few engineers, creating delays and forcing a trade-off between speed and quality. The team introduced Claude Code–based screening reviews, then expanded the approach into a broader system for improving PR creation, review accuracy, and continuous improvement. ## PR Review Bottlenecks - In late 2024, review responsibilities were concentrated on the tech lead and one other engineer. - Reviewers were simultaneously implementing features and reviewing code, causing PR queues to grow. - The main problems were: - Authors could not move to their next tasks while waiting for reviews. - Review work consumed most of the day. - Large PRs had to be reviewed quickly, increasing the risk of missed bugs. - This created a single point of failure and exposed the trade-off between thoroughness and development speed. - The launch of a dedicated frontend team in early 2025 provided an opportunity to redesign the review process. ## Introducing AI Screening Reviews - The team first tried having AI summarize PR changes before review. - Although summaries made changes easier to understand, AI did not sufficiently reduce the work of tracing dependencies or identifying hidden problems. - Manually pasting prompts for every review also made the approach inconvenient, so it was abandoned after about two weeks. - The introduction of Claude Code in summer 2025 changed the situation because reusable custom commands eliminated repetitive prompt preparation. - AI screening reviews now perform an initial inspection before a human reviewer makes the final judgment. - This changes the process from “humans inspect everything” to a two-stage model: - AI analyzes the PR, its impact, coding conventions, and possible risks. - A human reviewer validates the analysis and makes the final decision. ## Claude Code Custom Review Commands The custom command requests that Claude Code: - Summarize the PR and its affected areas. - Explain the before-and-after changes for each file. - Check coding and naming conventions. - Investigate dependent files and broader codebase impact. - Identify potential bugs, security issues, performance problems, code smells, and unintended side effects. - Suggest concise, respectful review comments for the author. - Classify comments with labels such as `[must]`, `[want]`, `[imo]`, `[ask]`, `[nits]`, and `[info]`. - Determine whether additional tests are needed based on existing project practices. The command uses GitHub CLI operations such as: - `gh pr view --json title,body,files,url` - `gh pr diff` - `gh pr view --comments` - GitHub API calls for line-level comments - `gh pr checkout` when the relevant branch is not currently checked out The review procedure is deliberately structured: 1. Confirm the review requirements. 2. Understand the PR’s overall purpose and background. 3. Review each changed file in detail. 4. Investigate dependencies across the codebase. 5. Produce a final assessment and suggested comments. The same screening process can help both reviewers and PR authors. Reviewers use it to reduce preparation time and understand impact, while authors can run it before requesting review to fix likely issues in advance. ## Expanding Beyond AI Screening After seeing benefits from screening reviews, the team created a broader improvement framework spanning technology and team culture. It was organized around four connected goals: - Improving efficiency. - Establishing a foundation for review accuracy. - Building review-oriented team culture. - Creating a mechanism for continuous improvement. The approach treats review optimization as an ongoing cycle rather than a one-time tool deployment. ## Automating PR Creation The team also uses AI to reduce the effort required to create PRs. - Git operations such as branch creation, commits, and PR creation are automated. - AI analyzes the commit diff to generate: - A PR title. - A summary of the changes. - Background and motivation. - Other required PR template fields. - Standardized and more complete PR descriptions provide better context for both human reviewers and AI screening. - Improving PR quality at the creation stage also increases the accuracy and consistency of later reviews. ## Practical Recommendation AI should support—not replace—reviewer judgment. Teams should begin by standardizing the review workflow, encode that workflow in reusable AI commands, and measure whether review time, PR waiting time, and review quality improve. Combining AI screening with better PR context, dependency analysis, clear comment conventions, and continuous process refinement offers a more sustainable solution than relying on individual reviewers.

Read original(opens in new tab)
line1 min readCurated summary

List of Articles for ‘Orchestration Development Workshop,’ an Internal Workshop to Enhance AI Utilization Skills

LY Corporation runs the “Orchestration Development Workshop” to help engineers apply AI more effectively in real development work. The workshop focuses on connecting multiple AI tools and agents to amplify creativity through collaborative learning and creation. A related blog series will share workshop topics, beginning with an example of using AI to reduce PR review delays and improve review culture. ## Orchestration Development Workshop - Targets engineers involved in development at LY Corporation. - Emphasizes practical, workplace-oriented AI skills rather than purely theoretical knowledge. - Uses “orchestrating multiple AIs to maximize creativity” as its central theme. - Builds on the idea that organizational learning is essential for successfully adopting AI. ## Blog Series - The series will publish workshop content incrementally. - The article list will be updated as new posts are released. - The listed update date is April 10, 2026. ## AI-Assisted PR Reviews - The first installment addresses bottlenecks and delays in pull-request reviews. - It explains how AI-assisted PR review support can help resolve review stagnation. - It also describes an internal workshop designed to change team review practices and culture. The series presents AI adoption as an organizational and collaborative practice, with PR review automation serving as an initial example of measurable workflow improvement.

Read original(opens in new tab)
meta3 min readCurated summary

Escaping the Fork: How Meta Modernized WebRTC Across 50+ Use Cases

Meta escaped the “forking trap” by replacing its divergent WebRTC fork with a modular architecture based on the latest upstream release. The system builds legacy and current WebRTC versions side by side, enabling runtime A/B testing across more than 50 use cases before rollout. This improved performance, binary size, and security while establishing a repeatable process for continuous upstream upgrades. ## Why the WebRTC Fork Became a Problem - Meta’s RTC stack supports Messenger, Instagram video calls, Cloud Gaming, and Meta Quest casting. - Internal optimizations and bug fixes gradually caused its WebRTC fork to diverge from upstream. - As the fork accumulated custom changes, merging community improvements became increasingly expensive and risky. - A one-time upgrade was impractical because WebRTC serves billions of users across diverse devices and environments. ## Requirements for a Sustainable Upgrade Strategy - Meta needed to: - Run legacy and upstream-based WebRTC implementations simultaneously. - Dynamically assign users to either version for safe A/B testing. - Statically link both versions into the same application. - Maintain custom patches in a monorepo without repeatedly rebuilding the migration process. - Standard patch-file workflows were considered difficult to scale for Meta’s large codebase. ## Shim Layer and Dual-Stack Architecture - A shim library was placed between application code and WebRTC. - Applications call a unified, version-neutral API rather than calling either WebRTC implementation directly. - A runtime “flavor” configuration routes each call to either the legacy or latest implementation. - Shimming at the lowest practical layer avoided duplicating the higher-level call orchestration library: - Full duplication would have added about 38 MB uncompressed. - The shim-based design added roughly 5 MB, an 87% reduction. ## Resolving C++ Symbol Collisions - Linking two WebRTC copies normally violates the C++ One Definition Rule and creates thousands of duplicate symbols. - Meta automated namespace rewriting: - `webrtc::` in the current version became `webrtc_latest::`. - The legacy version became `webrtc_legacy::`. - Global functions, variables, and classes outside namespaces were moved into namespaces where possible or assigned flavor-specific names. - Macro conflicts, including `RTC_CHECK` and `RTC_LOG`, were addressed by: - Removing unnecessary includes. - Renaming infrequently used macros. - Sharing modules such as `rtc_base` between versions to reduce duplication and shimming work. ## Preserving Backward Compatibility - Renaming symbols could have broken existing call sites, especially code built for only one WebRTC flavor. - An initial solution forward-declared every required symbol, but this created a large and fragile maintenance burden. - The improved approach used C++ `using` declarations to bulk-import a flavor namespace into the familiar `webrtc::` namespace. - This preserved existing source-level APIs without adding binary overhead, while allowing Meta to migrate selected call sites incrementally. ## Runtime Flavor Dispatch - Shim adapters and converters must instantiate objects from either the legacy or current namespace. - A template-based helper library keeps shared adapter logic in one place. - Template specializations handle version-specific behavior. - A global flavor enum, initialized during application startup, determines which WebRTC implementation is used. - The design also supports single-flavor builds during the transition. Meta’s approach demonstrates that large internal modifications do not have to require a permanent fork. A low-level shim, automated renamespacing, compatibility imports, and template-based dispatch provide a practical foundation for continuously rebasing custom functionality onto upstream WebRTC while safely validating each release through A/B testing.

Read original(opens in new tab)
figma2 min readCurated summary

Turning Prompts into Five Scalable Workflows with Figma Weave | Figma Blog

Figma Weave presents AI creation as a scalable, editable workflow rather than a one-off prompt. Its canvas connects AI models and processing nodes so creators can branch, refine, and reuse each step while maintaining control over imagery, video, audio, and 3D output. The article introduces five workflows, beginning with a method for deriving a reusable visual style from multiple reference images. ## Figma Weave as a Creative Workflow Canvas - Figma Weave evolved from Weavy, which Figma acquired to expand its capabilities in: - Image and video generation - Animation and motion design - Audio and 3D creation - VFX and professional editing - Users can chain prompts and AI nodes together, moving from references to finished assets without losing the ability to revise intermediate steps. - Figma has published more than 20 Community templates covering tasks such as: - Turning images into videos - Generating 3D models - Combining visual references - Comparing image-generation models ## Why Workflows Are More Scalable Than Single Prompts - A single prompt produces one interpretation of a style. - A workflow lets creators independently adjust how strongly each reference influences the result. - Individual stages can be reshaped, reused, and applied across multiple assets and channels. - The example brand, Epoch, demonstrates how the system can support a consistent visual identity based on distorted textures and 3D natural forms. ## Combining Two Images into a Reusable Style Guide - The first workflow combines a hibiscus flower and a rock face from Epoch’s existing visual references. - Each image is processed through an **Image Describer node**, which extracts attributes such as: - Texture - Color - Lighting - Composition - The resulting text descriptions can be edited and merged into a new style definition. - The balance between the two references can be adjusted until the desired blend is achieved. - The combined style can then be tested across different image-generation models, helping the team validate the look at scale. - The output is treated as a reusable style system rather than a single prompt for one image. The practical recommendation is to build visual direction as a modular workflow: analyze existing references, combine and tune their characteristics, and preserve the resulting style definition for reuse in future assets.

Read original(opens in new tab)
gitlab2 min readCurated summary

5 ways GitLab pipeline logic solves engineering problems

GitLab’s pipeline model addresses complex CI/CD needs by combining composable features rather than relying on a single linear workflow. Parent-child pipelines, DAG execution, and multi-project triggers help teams scale monorepos and coordinate services across repositories while preserving clear ownership and failure visibility. The article argues that these patterns make pipelines both faster and easier to maintain. ## Monorepos: Parent-child pipelines and DAG execution - A monorepo containing frontend, backend, and documentation projects should not rebuild everything for every change. - Parent pipelines can trigger child pipelines for individual services using `trigger: include`. - Multiple included files are merged into one child pipeline, allowing jobs across files to share context and reference one another with `needs:`. - `strategy: depend` makes the parent wait for child pipelines and report one overall success or failure while retaining detailed drill-down. - Each service can own its pipeline configuration, reducing the risk that changes in one service break another. - DAG execution with `needs:` allows dependent jobs to start as soon as their prerequisites finish instead of waiting for an entire stage. - For example, API tests can begin immediately after the API build completes, without waiting for unrelated jobs. ## Microservices: Cross-repository pipelines - When frontend and backend services live in separate repositories, independent pipelines may miss integration failures. - GitLab multi-project pipelines allow one repository to trigger and await a pipeline in another project. - The frontend can generate an API contract artifact, publish it, and trigger the backend pipeline with `strategy: depend`. - The backend downloads the artifact through the GitLab Jobs API using `CI_JOB_TOKEN`. - An integration test can reject breaking API changes and propagate the failure back to the frontend pipeline. - The backend job uses `CI_PIPELINE_SOURCE == "pipeline"` so the contract validation runs only when initiated by the frontend, not during ordinary backend pushes. - The frontend project identifier is supplied through a CI/CD variable such as `FRONTEND_PROJECT_ID`. These patterns let teams reduce unnecessary work, preserve service-level ownership, and make cross-service compatibility checks part of the delivery process.

Read original(opens in new tab)
toss3 min readCurated summary

It Almost Ended Up Ugly - The Making of Toss Front 2

Toss redesigned its Front 2 payment terminal by addressing real-world usability problems rather than settling for technically workable solutions. The redesign moved NFC to the front, made the card reader field-replaceable, and reworked the internal structure to simplify removal. The result was a smaller, cleaner device that improved both customer experience and repairability. ## Moving NFC to the Front - The first-generation terminal placed NFC on the right side because other components interfered with the signal. - This was inconvenient in narrow retail spaces, where users had little room to tap cards or phones. - Several alternatives were tested: - Moving the card reader upward made card insertion awkward and strained users’ wrists. - Enlarging the top made the vertically oriented terminal look excessively long and displaced the camera, complicating barcode and face-payment use. - Placing NFC around the camera caused interference, producing camera shake and reducing recognition accuracy. - The team reframed the problem by asking whether the display’s metal backing could be replaced. - A customized plastic backing allowed NFC signals to pass through while reinforced glass preserved the display’s rigidity. - Despite higher manufacturing complexity and cost, the design was successfully mass-produced, enabling reliable front-facing NFC without sacrificing appearance. ## Designing a Replaceable Card Reader - The first-generation card reader was integrated into the main body, so failures required repairing or replacing the entire terminal. - Repairs took more than a week on average, forcing stores to use backup devices and distributors to maintain extra inventory. - Front 2 introduced a docked, replaceable card reader designed for easy on-site replacement. - A USB-C connector was selected because users already understand how to connect and disconnect it. - The connector provided stable attachment without exposing additional brackets or mechanisms, preserving the product’s clean appearance. ## Making Removal Simple - Once the reader used USB-C, the team needed a way to remove it without adding buttons, levers, or protruding parts. - The simplest approach was to insert the reader from the front and push it out from behind, but existing power and network connections blocked the necessary space. - Instead of adding another mechanism, the team redesigned the internal layout from scratch. - The circuit board and connectors were tilted toward the top of the device, requiring redesigned and inverted cable components. - This created enough room to remove the reader without tools. - The revised layout also made the cables easier to see and connect. ## Results of Front 2 - Front-facing NFC made payments more natural on crowded counters. - The USB-C dock reduced the difficulty of replacing a failed card reader. - The terminal became smaller while retaining its visual simplicity and design quality. - Front 2 surpassed the first generation’s sales immediately after launch, while previous customer complaints shifted toward positive feedback. ## Building Quality Through Persistent Questions - The article argues that product quality comes from repeatedly challenging an acceptable technical solution. - The team focused on finding answers that were right for users and the overall product experience, not merely answers that functioned. - For similar design problems, the recommended questions are: - Can everyone use the design easily, including in edge cases? - What is the fundamental problem? - Is the current solution truly the best one? - If not, can the problem and solution direction be redefined? The practical lesson is to keep revisiting the problem until the solution is not only feasible, but genuinely appropriate for users.

Read original(opens in new tab)
google3 min readCurated summary

ConvApparel: Measuring and bridging the realism gap in user simulators

ConvApparel addresses the “realism gap” between LLM-based user simulators and genuine human behavior. It combines over 4,000 human-AI shopping conversations with a controlled Good-versus-Bad agent setup and evaluates simulators through statistical alignment, human-likeness, and counterfactual adaptation. The framework aims to determine whether simulators genuinely model human reactions or merely reproduce patterns from their training data. ## Why User Simulator Realism Matters - Conversational agents often fail during long, multi-turn interactions by forgetting constraints or producing irrelevant responses. - Human testing provides valuable feedback but is expensive, slow, and difficult to scale. - LLM-based user simulators offer a scalable alternative, but often behave unlike real users: - They may be excessively verbose. - They can lack consistent personas or coherent preferences. - They may possess unrealistic, encyclopedic knowledge. - They are often unusually patient and assistant-like. - Training systems only against unrealistic simulators may cause them to fail with real users. ## The Need for Counterfactual Validation - A simulator should respond plausibly not only to situations represented in its training data, but also to novel assistant behaviors. - The authors introduce **counterfactual validation**: training a simulator on helpful-agent conversations, then testing it against an unexpectedly frustrating agent. - A realistic simulator should recognize poor assistance and show increased frustration, reduced satisfaction, and behavior changes similar to those of real users. - This tests whether the simulator has learned general human behavior rather than memorized training patterns. ## The ConvApparel Dataset - ConvApparel contains more than 4,000 human-AI multi-turn conversations and nearly 15,000 total turns in the apparel-shopping domain. - Participants were unknowingly assigned to one of two recommendation agents: - **Good agent:** Helpful, efficient, and supported by robust search. - **Bad agent:** Intentionally confusing, tangential, and based on degraded search retrieval. - The dataset captures reactions ranging from satisfaction to significant annoyance. - Participants provided turn-by-turn retrospective annotations, including: - Satisfaction - Frustration - Likelihood of making a purchase ## Three-Part Evaluation Framework ### Population-Level Statistical Alignment - Simulated conversations are compared with human conversations using aggregate measures such as: - Conversation length - Words per turn - Dialogue acts, including rejecting recommendations - This reveals whether simulators reproduce broad behavioral distributions. ### Human-Likeness Score - An automated discriminator is trained on human and simulated conversations. - It produces a probability indicating how human-like a conversation appears. - The score is intended to detect subtle stylistic differences that simple statistics may miss. ### Counterfactual Validation - A simulator is trained only on conversations with the Good agent. - It then interacts with the unseen Bad agent. - High-fidelity simulation should produce a human-like increase in frustration and decline in satisfaction when the assistant behaves poorly. ## Simulator Configurations The experiments compare three Gemini-based user simulators: - **Prompted simulator:** Uses high-level behavioral instructions without additional task-specific training. - **In-context learning (ICL) simulator:** Retrieves semantically similar human conversations from ConvApparel and supplies them as examples at each turn. - **Supervised fine-tuning (SFT) simulator:** Trains a Gemini 2.5 Flash model directly on the dataset. The post presents ConvApparel as a structured way to measure simulator realism and test whether simulated users can adapt to assistant behavior outside their training distribution. Its central recommendation is to evaluate user simulators not only by surface-level similarity, but also by how naturally they react to unexpected failures.

Read original(opens in new tab)
meta2 min readCurated summary

Trust But Canary: Configuration Safety at Scale

As AI accelerates software development, stronger safeguards are needed to prevent faster mistakes from becoming larger incidents. Meta’s Configurations team uses canarying, progressive rollouts, health checks, and monitoring to detect regressions early. Data and AI also help reduce alert noise and speed up identifying the changes responsible for failures. ## Safe Configuration Rollouts - Meta deploys configuration changes gradually rather than releasing them everywhere at once. - Canarying exposes changes to a small subset of systems or users first. - Progressive rollouts expand the deployment only when monitoring indicates that the change is healthy. - These practices limit the impact of faulty configurations and provide opportunities to stop or reverse a rollout. ## Monitoring and Health Checks - Automated health checks and operational signals help identify regressions soon after deployment. - Monitoring provides evidence for deciding whether a rollout should continue, pause, or be rolled back. - Early detection is especially important at Meta’s scale, where a small configuration error can affect many systems. ## Learning from Incidents - Incident reviews focus on improving tools, processes, and safeguards rather than assigning blame to individuals. - The goal is to make future failures less likely and reduce their potential impact. - These reviews turn operational problems into improvements across the configuration management system. ## AI-Assisted Operations - Data-driven techniques reduce alert noise so engineers can focus on meaningful signals. - AI and machine learning help speed up bisection, narrowing down which change introduced a problem. - Faster diagnosis can shorten recovery times and make progressive deployment practices more effective. The episode recommends combining gradual releases, strong observability, blameless incident reviews, and AI-assisted analysis to keep increasingly rapid development safe at scale.

Read original(opens in new tab)
cloudflare3 min readCurated summary

From bytecode to bytes- automated magic packet generation

Classic BPF filters can hide Linux malware until a precisely crafted “magic” packet arrives, but manually reverse-engineering large filters is slow and error-prone. The post presents a symbolic-execution approach using the Z3 theorem prover to model BPF instructions as packet constraints and automatically generate triggering packets. This reduces analysis from hours of manual work to seconds, even for filters exceeding 100 instructions. ## Why BPF Filters Are Difficult to Analyze - Classic BPF is a small, efficient virtual machine used to filter network traffic inside the Linux kernel. - Unlike eBPF, classic BPF has a simple two-register design but can still contain many conditional jumps and packet-offset calculations. - Malware authors exploit BPF because kernel-level filtering can hide traffic from ordinary user-space monitoring tools. - Short programs are manageable manually, but complexity grows rapidly as filters reach 100 or more instructions. - The core problem is determining which packet bytes satisfy the conditions along an accepting execution path. ## BPFDoor as a Practical Example - BPFDoor is a stealthy Linux backdoor associated with cyberespionage campaigns and groups including Red Menshen/Earth Bluecrow. - It uses BPF to inspect incoming traffic without listening on a dedicated open port. - The example filter checks: - IPv6 or IPv4 EtherType. - UDP protocol. - DNS destination port 53. - Fragmentation status for IPv4 packets. - The IPv4 header length when locating the UDP destination port. - The filter contains two paths leading to acceptance: - An IPv6 UDP packet destined for port 53. - A non-fragmented IPv4 UDP packet destined for port 53. - These paths expose the byte offsets and values that a generated packet must satisfy. ## Finding the Shortest Accepting Path - The tool explores the BPF control-flow graph using a queue. - Each queued item records: - The next instruction pointer. - The sequence of instructions already traversed. - Conditional jumps are explored in both directions. - Paths ending in a drop result are discarded, while paths reaching a nonzero return value are recorded as accepting paths. - Breadth-first traversal prioritizes paths with fewer conditions, helping identify the shortest route to acceptance. - Unconditional jumps are followed directly, while conditional branches enqueue true and false destinations in order of path length. ## Turning Paths into Packets - Once accepting paths are identified, each branch condition becomes a constraint on packet contents. - The required byte offsets, widths, and values can be collected from the executed instructions. - Symbolic execution represents these checks as constraints instead of requiring analysts to reason through every instruction manually. - Z3 can then solve the resulting constraint set and produce packet bytes that satisfy the selected accepting path. - This approach is especially useful for large or heavily branched BPF programs where manual packet construction becomes impractical. The recommended workflow is to combine control-flow exploration with symbolic constraint solving: first identify viable accepting paths, then use Z3 to generate packets satisfying their byte-level requirements. This automates a formerly labor-intensive part of malware analysis and makes complex BPF-based backdoors much faster to investigate.

Read original(opens in new tab)
line3 min readCurated summary

The Key to AI Utilization Lies in 'Organizational Learning' - The Start of the Orchestration Development Workshop

LY Corporation is moving from simply adopting AI tools to building with AI as a collaborative development partner. Its new Orchestration Development Workshop teaches engineers to coordinate multiple AI systems across coding, testing, reviews, incident analysis, and other workflows. The initiative aims not only to improve efficiency but to free engineers from repetitive work so they can focus on more creative, high-value challenges. ## From AI Adoption to AI Collaboration - AI-assisted development and operations are spreading rapidly across LY Corporation. - Engineers use generative AI for code generation and testing, while combining it with non-generative AI for analysis and operational optimization. - Despite broader adoption, employees differ significantly in how deeply they use AI in their daily work. - The workshop was created to help the organization evolve from “using AI” to “creating alongside AI.” ## Orchestration: Coordinating Multiple AI Systems - “Orchestration” refers to combining multiple AIs, along with human input, to produce a complete outcome. - Example workflows include: - Generating code automatically from a Jira ticket. - Having AI run tests, conduct reviews, and create a pull request. - Analyzing a Slack incident report, estimating the cause, and proposing a fix. - The workshop turns these emerging practices into hands-on learning rather than passive demonstrations. ## A Hands-On, Interactive Learning Model - Participants follow instructors in real time and perform the same tasks themselves. - Zoom conversations and Slack questions create two-way communication during the session. - Instructors and representative participants explore solutions to problems as they arise. - The goal is for attendees to gain skills they can reproduce in their own projects, not merely acquire theoretical knowledge. ## Organization-Wide Support Through Guilds and DevRel - The initiative is designed to avoid depending on individual enthusiasm. - Three complementary functions support continuous growth: - **DevRel:** Drives the program and promotes adoption. - **Guilds:** Contribute practical insights from engineering teams. - **TD:** Helps maintain quality and reproducibility. - This structure supports consistent content quality and enables AI knowledge to spread across the company. ## Beyond Efficiency: Unlocking Engineering Creativity - LY Corporation views AI as more than a way to complete tasks faster. - By delegating repetitive work to AI, engineers can spend more time on creative and strategically valuable activities. - The organization aims to move beyond a model where AI writes code and humans only review it. - Instead, engineers should collaborate with AI from the design stage through implementation. ## Future Direction - LY Corporation plans to share lessons from the workshops through external channels such as its technology blog. - Future topics will include both generative and non-generative AI. - The broader goal is to provide practical guidance for engineers building new workflows with AI. The workshop represents a structured way to turn AI experimentation into repeatable organizational practice, helping engineers coordinate multiple AI tools while preserving human creativity and judgment.

Read original(opens in new tab)
google3 min readCurated summary

Improving the academic workflow: Introducing two AI agents for better figures and peer review

AI is being positioned as an active participant in academic research, not merely a tool for drafting text. The post introduces PaperVizAgent, which creates publication-ready figures, and ScholarPeer, which produces literature-grounded peer reviews. Both use multi-agent workflows and iterative verification to reduce researchers’ administrative burden while improving visual quality and review rigor. ## PaperVizAgent: Generating Publication-Ready Figures - PaperVizAgent converts manuscript text and a detailed figure caption into academic illustrations. - It uses five specialized agents: - **Retriever:** Finds relevant literature and reference figures. - **Planner:** Organizes the technical content. - **Stylist:** Develops appropriate visual and aesthetic guidelines. - **Visualizer:** Produces images or executable Python code for statistical plots. - **Critic:** Checks the result against the source text and requests revisions. - The critic-driven refinement loop is designed to ensure that figures are both technically faithful and visually clear. - Inputs typically include: - The manuscript’s method or technical sections. - A communicative-intent description explaining what the figure should convey. ### Evaluation Results - PaperVizAgent was compared with direct prompting, few-shot prompting, GPT-Image-1.5, Nano-Banana-Pro, and Paper2Any. - Figures were scored from 0 to 100 on: - Faithfulness - Conciseness - Readability - Aesthetics - It achieved an overall score of **60.2**, exceeding the human baseline of **50.0** and outperforming the evaluated automated systems. - Its strongest results were in conciseness and aesthetics, while its statistical plots reached human-competitive quality. ## ScholarPeer: Automating Rigorous Peer Review - ScholarPeer is a search-enabled, context-aware multi-agent system designed to emulate the workflow of a senior academic reviewer. - Rather than treating review as simple text generation, it combines literature retrieval, adversarial checking, and technical verification. - Its main components include: - A **sub-domain historian** that builds a current domain narrative from literature. - A **baseline scout** that searches for overlooked datasets, methods, and comparisons. - A **multi-aspect Q&A engine** that tests novelty and technical claims. - A **review generator** that follows conference-specific review guidelines. - The resulting review includes a summary, strengths, weaknesses, and questions for the authors. ### Evaluation Results - ScholarPeer was evaluated on public datasets against fine-tuned models and other agentic reviewing systems. - Its active web-search and verification process produced highly critical reviews grounded in existing research. - Side-by-side evaluations showed strong win rates against competing automated reviewers. - The system also narrowed the gap between AI-generated reviews and human reviews in terms of realism, diversity, and alignment with expert judgments. ## Implications for Academic Research - The two agents address separate bottlenecks in the publication process: - PaperVizAgent improves technical communication through better figures. - ScholarPeer helps scale peer review amid growing submission volumes and reviewer fatigue. - Their multi-agent designs suggest that specialized agents, coordinated through retrieval and iterative critique, may be more effective than a single general-purpose language model. - The systems are intended to support researchers rather than replace scientific judgment. Researchers could use PaperVizAgent for early figure prototyping and ScholarPeer for preliminary, literature-informed critique, while retaining human oversight for final scientific and editorial decisions.

Read original(opens in new tab)