Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

Powering the agents: Workers AI now runs large models, starting with Kimi K2.5

Cloudflare is expanding Workers AI beyond smaller models by adding Moonshot AI’s Kimi K2.5, a frontier open-source model designed for agentic workloads. With a 256k context window, tool calling, vision, and structured outputs, Kimi can power an agent’s full lifecycle directly on Cloudflare’s platform. Cloudflare argues that its price-performance makes open-source models essential as personal and enterprise agents dramatically increase inference demand. ## Kimi K2.5’s Price-Performance Advantage - Cloudflare uses Kimi internally for: - Agentic coding through OpenCode - Automated code review via the Bonk public code review agent - Security analysis of Cloudflare codebases - A security-review agent processes more than 7 billion tokens daily and has found over 15 confirmed issues in one codebase. - Compared with a mid-tier proprietary model, switching to Kimi reduced the estimated cost of this workload by 77%, from roughly $2.4 million annually. - As employees increasingly run multiple agents continuously, inference costs become a major barrier to scaling. - Cloudflare positions open-source, frontier-quality models as a more economical alternative to proprietary systems. ## Serving Large Models on Workers AI - Supporting Kimi required upgrades to Workers AI’s inference stack, which historically focused on smaller models. - Cloudflare uses its proprietary Infire inference engine and custom kernels to improve: - Model performance - GPU utilization - Throughput - The platform applies advanced serving strategies such as: - Data, tensor, and expert parallelization - Disaggregated prefill, separating input processing from generation across machines - Workers AI handles these infrastructure optimizations so developers do not need specialized machine learning, DevOps, or reliability engineering expertise. ## Prefix Caching for Agent Workloads - Agents frequently resend large prompts containing: - System instructions - Tool definitions - MCP server tools - Conversation history - Entire codebases - Prefix caching avoids reprocessing unchanged input tokens during multi-turn interactions. - This reduces prefill work, improving: - Time to First Token (TTFT) - Tokens Per Second (TPS) - Overall inference cost - Workers AI now exposes cached tokens as a usage metric and charges less for them than regular input tokens. - Cloudflare has also introduced techniques to improve cache hit rates. ## Session Affinity - Workers AI provides an `x-session-affinity` header to route requests from the same session or agent to the same model instance. - Keeping requests on the same instance increases prefix-cache reuse. - Higher cache hit rates lead to faster responses, greater throughput, and lower costs. - Clients should provide a unique session or agent identifier with the header. Cloudflare’s recommendation is to use Workers AI when building agents that need frontier-level reasoning without the cost and operational burden of proprietary models or self-hosted infrastructure.

Read original(opens in new tab)
github3 min readCurated summary

How Squad runs coordinated AI agents inside your repository

Squad is an open-source GitHub Copilot project that places a preconfigured team of AI agents directly inside a repository. Rather than relying on a single chatbot or complex orchestration infrastructure, it coordinates specialized agents for design, implementation, testing, documentation, and review. Its core argument is that repository-native, versioned context makes multi-agent development more accessible, inspectable, and resilient. ## Coordinating Specialized Agents - Install Squad with `npm install -g @bradygaster/squad-cli`, then run `squad init` in a repository. - The setup creates roles such as lead, frontend developer, backend developer, tester, and documentation specialist. - A coordinator interprets natural-language requests, loads repository context, and assigns work to specialists. - Agents can work in parallel, create files and branches, write tests, and open pull requests. - They use shared decisions and project history rather than requiring every detail to be repeated in prompts. - Testing and review happen within the workflow: - Testers evaluate implementations and reject failing code. - A rejected author is prevented from revising its own work. - Another agent must address the problems, providing a more independent review. - Developers still answer questions, correct assumptions, and review and merge pull requests; Squad is collaborative orchestration rather than full autonomy. ## Repository-Based Shared Memory - Squad uses a “drop-box” model instead of depending on live chat synchronization or complex vector databases. - Architectural decisions, library choices, and conventions are appended to a versioned `decisions.md` file. - This creates: - Persistent shared knowledge - An understandable audit trail - Recovery after disconnects or restarts - Memory that can be reviewed and changed like code ## Replicating Context Across Agents - The coordinator remains a thin router instead of attempting to manage all implementation work. - Each specialist runs in its own inference call with an independent context window. - This replicates relevant repository context across agents rather than splitting one limited context among multiple roles. - Parallel, independent contexts reduce the risk that project-management instructions and other agents’ reasoning crowd out the actual coding task. - Supported models may provide context windows of up to 200,000 tokens. ## Versioned Agent Identities and History - Each agent’s behavior is primarily defined by repository files: - A charter describing its role and responsibilities - A history recording previous work - Shared team decisions - These files live in `.squad/` alongside the application code. - Cloning a repository also restores the team’s accumulated knowledge, making the agents effectively pre-onboarded. - Keeping memory in plain text makes it inspectable, versioned, and independent of hidden model state. ## Lowering the Barrier to Multi-Agent Development Squad’s main goal is to make agentic workflows practical without requiring users to build orchestration layers, configure databases, or master advanced prompt engineering. Its repository-native design favors simple setup, transparent memory, independent review, and recoverable project context. Developers interested in this approach can install Squad and experiment with it directly in the project repository.

Read original(opens in new tab)
aws4 min readCurated summary

20 years in the AWS Cloud – how time flies! | Amazon Web Services

AWS’s 20-year evolution reflects a shift from foundational cloud infrastructure to managed services for AI, automation, and agentic applications. The author argues that AWS’s most important innovations come from responding to customer needs rather than chasing every fashionable technology. Personal experiences with AWS and its community illustrate how cloud services have enabled developers, researchers, and businesses to pursue previously impractical projects. ## AWS’s Impact on the Author’s Career - The author met AWS blogger Jeff Barr in Seoul in 2006, shortly after Amazon began promoting API-based services. - Inspired by Barr, the author began building APIs for third-party developers and later used AWS for large-scale academic research. - The author’s company became one of Korea’s earliest AWS customers in 2014. - AWS helped make advanced computing capabilities accessible to individuals, startups, researchers, and enterprises. ## Innovation Driven by Customer Needs - AWS has grown to more than 240 cloud services and launches thousands of features each year. - The author highlights the importance of distinguishing genuine technological trends from temporary distractions. - AWS’s evolution spans deep learning, generative AI based on large language models, and today’s agentic AI. - The central innovation principle is to listen to customers and solve their most important problems, rather than adopting technology simply because it is fashionable. ## Major AWS Milestones The article recalls foundational services from AWS’s first decade, including: - Amazon S3 and EC2 in 2006 - Amazon RDS and VPC in 2009 - DynamoDB and Redshift in 2012 - WorkSpaces and Kinesis in 2013 - AWS Lambda in 2014 - AWS IoT in 2015 ## Containers and Serverless Databases - Amazon ECS, launched in 2014, simplified running containers across managed EC2 clusters. - Amazon EKS later added managed Kubernetes, while AWS Fargate enabled serverless container deployment. - Amazon Aurora provided highly available relational databases at scale. - Aurora Serverless evolved from version 1 to version 2, which can scale down to zero. - Aurora DSQL, launched in 2025, extends the serverless model to distributed SQL workloads requiring continuous availability. ## Making Machine Learning More Accessible - Amazon SageMaker, launched in 2017, provided an end-to-end managed environment for building, training, and deploying ML models. - In 2024, AWS introduced the next-generation SageMaker platform for data, analytics, and AI, along with SageMaker AI for model development and deployment. - AWS also developed specialized hardware: - Inferentia for low-latency inference - Trainium for high-performance AI training - Trainium3 UltraServers for improved economics in generative AI workloads ## Improving Cloud Price Performance - EC2 A1 instances introduced AWS Graviton processors based on Arm architecture. - Later Graviton generations expanded price-performance benefits across services such as ECS, EKS, Lambda, RDS, ElastiCache, EMR, and OpenSearch Service. - More than 90,000 customers have reportedly adopted Graviton-based infrastructure. ## Hybrid Cloud and Edge Computing - AWS Outposts brings AWS infrastructure and services into customer data centers and edge locations. - Available configurations range from 1U and 2U servers to 42U racks and multi-rack deployments. - Customers use Outposts for low-latency access, local processing, data residency, and applications with on-premises dependencies. ## Generative AI and Agentic Development - Amazon Bedrock provides access to multiple AI models and managed capabilities for building secure generative AI applications. - Bedrock AgentCore extends the platform to deploying and operating agents at scale. - More than 100,000 customers use Bedrock for personalization, workflow automation, and insight generation. - Amazon CodeWhisperer evolved into Amazon Q Developer, adding conversational assistance, project-based generation, and code transformation. - The service later evolved into Kiro, an agentic development tool centered on spec-driven development and autonomous coding tasks. - AWS expanded model choice through Amazon Titan and Amazon Nova, including services for building frontier models and browser-automation agents. AWS’s history suggests that the strongest path forward is to use AI and cloud services to address concrete customer and business challenges. The author’s examples present AWS as an evolving platform whose value comes not only from individual launches, but from steadily making advanced infrastructure, machine learning, and autonomous software development more accessible.

Read original(opens in new tab)
slack2 min readCurated summary

How Slack Rebuilt Notifications 📣

Slack rebuilt its notification system to reduce noise by replacing years of inconsistent, tightly coupled behavior with a unified model. The redesign separates what activity users see from how they receive interruptions, while aligning desktop and mobile settings. By combining backend migration strategies, auto-saving controls, and shared UI patterns, Slack aims to make notifications predictable and easier to manage. ## Diagnosing Notification Overload - Notification frustration is common, especially for users in many channels. - Notification issues are among Slack’s top three sources of Customer Experience tickets. - The underlying problem was architectural as well as behavioral: - Desktop and mobile used conflicting preference systems. - Equivalent settings, such as “Nothing” and “Off,” behaved differently across clients. - Activity preferences were coupled to push delivery. - Settings could fall out of sync between desktop and mobile. - Advanced options were scattered or difficult to discover. ## A Unified Notification Model Slack introduced a simpler set of controls: - Channel notifications now offer: - **All new posts** - **Mentions** - **Mute** - Push notifications have separate on/off controls across desktop and mobile. - Advanced features, including mobile “badge all unreads,” are easier to find. - Global preferences use consistent structure and language. - Simplified preference logic improves synchronization between clients. ## Refactoring Preferences Safely - Slack migrated users from four conflicting preference systems to a unified model. - The new model separates: - Desktop activity: **Everything** or **Mentions** - Desktop push: `desktop_push_enabled` set to `true` or `false` - Mobile activity and push behavior: **Everything**, **Mentions**, or **Nothing** - Rather than changing millions of database records directly, Slack used read-time interpretation to preserve backward compatibility and allow rollback. - Existing “Off” settings now behave as “Mentions” with push disabled. - A backfill populated the new desktop push preference based on users’ previous settings. - This preserves in-app awareness while allowing push interruptions to be controlled independently. ## Auto-Saving and Clearer Controls - The previous modal required users to press **Save**, which caused accidental abandoned changes. - The redesigned interface applies changes immediately through auto-save. - Users can independently choose what activity to see and how they want to receive it. - Shared React components replaced legacy mobile-specific UI code, improving consistency across platforms. - Users can now, for example, view all activity while receiving push notifications only for mentions. Slack’s approach demonstrates that reducing notification noise requires more than a visual redesign. Separating activity from delivery, simplifying preference states, and keeping clients synchronized gives users clearer and more reliable control over interruptions.

Read original(opens in new tab)
stripe3 min readCurated summary

Testing the impact of Adaptive Pricing across 1.5M subscription checkout sessions

Subscription businesses face major operational and payment challenges when pricing internationally, especially because exchange rates fluctuate across recurring billing cycles. Stripe’s Adaptive Pricing addresses this by displaying and charging customers in local currencies while managing FX conversion and renewal stability. An analysis of 1.5 million checkout sessions found that it increased signup conversion, payment authorization, and subscription lifetime value. ## The Challenge of Localized Subscription Pricing - Localizing prices requires businesses to manage: - FX risk and conversion fees - Currency-specific price lists - Ongoing accounting, reconciliation, and reporting - Subscriptions are more complex than one-time purchases because prices must remain predictable at every renewal. - Cross-border charges are more likely to fail than local-currency transactions. - In 2025, 80% of subscription transactions were still priced in a company’s default currency. ## How Adaptive Pricing Works - Stripe’s Adaptive Pricing, available through the Optimized Checkout Suite, automatically displays prices in a customer’s local currency. - Stripe handles currency conversion and related operational work. - A stability buffer helps keep renewal amounts consistent despite exchange-rate changes. - For example, a Brazilian customer might continue paying R$49.60 per month rather than seeing a different converted amount each billing cycle. - Significant exchange-rate movements can still cause a renewal amount to change, similar to adjustments made by card issuers. ## Measured Signup Improvements Stripe analyzed 1.5 million subscription checkout sessions, comparing Adaptive Pricing with a randomized 1% holdback group. - Average conversion increased by 4.7%. - Average payment authorization increased by 1.9%. - Lifetime value per checkout session increased by 5.4%. - Some businesses saw LTV gains above 30%. - Runway reported a 14% increase in LTV per session and 17.7% more LTV per subscription. Local prices reduce the need for customers to mentally convert costs and make recurring commitments feel more transparent. Charging locally can also improve authorization rates because cross-border payments are more likely to be declined. ## Retention and Long-Term Value - Customers who paid in local currency showed consistently higher retention than those charged in a company’s default currency. - Better signup conversion combined with more successful renewals increases the value generated by each checkout session. - Even small improvements in initial conversion and payment approval can compound into meaningful gains in subscription LTV. ## Scaling Globally - Businesses that do not localize subscription prices may be losing international revenue. - Adaptive Pricing lets companies offer local-currency pricing without building their own FX systems, currency-specific price catalogs, or renewal logic. - Stripe reports that more than 500,000 businesses, including over 16,000 subscription companies such as Cursor, Perplexity, and Runway, use Adaptive Pricing. For subscription companies expanding internationally, localized pricing can improve both initial purchase performance and long-term subscriber value while reducing the complexity of managing global billing.

Read original(opens in new tab)
line4 min readCurated summary

Utilizing SLI/SLO to Improve Reliability Part 1: SLI/SLO Framework and the Development Story of Service Status Check Tool LINE Status

Repeated SLI/SLO adoption revealed a common process that could be standardized across services. The team turned that process into a reusable framework and built “LINE Status,” an internal tool that automatically presents service health according to user experience rather than raw alerts. Together, these initiatives create a shared organizational language for understanding reliability and its impact on users. ## A Reusable SLI/SLO Framework After applying SLI/SLOs to several platforms and services, the SRE team identified recurring patterns independent of service type. They organized these patterns into a five-stage framework: - **Select critical user journeys (CUJs) and define SLIs** - Identify the experiences most important to users. - Define measurable SLIs that represent those experiences. - **Design instrumentation and metrics** - Build or adapt metrics suitable for each CUJ. - Use standardized naming based on Prometheus or OpenTelemetry. - **Create dashboards and recording rules** - Provide Grafana dashboards for quickly assessing SLO achievement. - Precompute complex PromQL operations to improve query performance. - **Set SLOs and alerts** - Begin with flexible targets, such as 99.9% availability over a 28-day rolling window, allowing roughly 40 minutes of downtime. - Define runbooks for responding to alerts. - Refine targets after operational data and experience accumulate. - **Establish error-budget governance** - Balance release speed against reliability. - Review objectives monthly or quarterly. - Adjust SLOs and processes as needed. The framework is currently distributed as a Confluence template containing guidance and FAQs, reducing the communication effort required from SREs during initial adoption. ## Moving from Alerts to User-Centered Service Status As more services adopted SLI/SLOs, the team wanted a consistent way to understand the health of services they did not directly operate. - The existing public LINE Status API page focused on external users and was updated manually during major incidents. - The new internal tool was intended to: - Represent the status of individual service components. - Update automatically from SLI/SLO alerts and outage data. - Show whether user experience was being affected. - Rather than simply reflecting whether an alert or outage existed, status was based on CUJ-related SLI performance and SLO achievement. - Only representative, high-value CUJs were exposed, avoiding unnecessary technical detail. ## LINE Status Architecture and Interface LINE Status was designed as more than an alert list. It collects events through webhooks, stores them in a separate database, and uses that data to track both current status and historical changes. - Technical SLI/SLO terms are translated into user-facing functions such as “Message Sending” or “Read Receipts.” - Status colors provide an immediate overview: - Green: normal - Yellow: event detected - Red: outage - The main page provides: - An overview of all services. - CUJ status within each service card. - AI-generated one-line summaries. - Service detail pages provide: - Recently affected items near the top. - Timeline-based event displays. - Monthly historical events. - The history page shows: - The scope of impact for each service during an event. - Past events organized by month. The initial implementation took about a month and was refined through colleague feedback. The author also used AI-assisted “vibe coding” for the frontend, emphasizing that clear, detailed requirements were more important than the development tool itself. ## Connecting the Framework and LINE Status Once a service adopts SLI/SLOs through the framework, it can be registered in LINE Status. This connects the definition of reliability objectives with an organization-wide view of service health. - Developers and operators can use the same CUJ-based standards. - Teams can focus on whether users are affected instead of interpreting isolated alerts. - During incidents, the tool helps identify impacted experiences quickly. - Over time, the approach may improve decision-making speed and cross-team communication. The team plans to refine CUJs, SLIs, and status-transition rules through continued operational experience. The practical goal is to make SLI/SLOs a common language for describing service health, enabling reliability practices to scale without depending heavily on individual teams or specialists.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab 18.10 brings AI-native triage and remediation

GitLab 18.10 adds AI-powered security features designed to reduce vulnerability triage noise and speed remediation. GitLab Duo Agent Platform can assess whether SAST and secret-detection findings are likely false positives, explain its reasoning, and— for verified SAST issues—generate tested fixes in merge requests. These capabilities are available to GitLab Ultimate customers using the Duo Agent Platform, with some features still in beta. ## SAST False Positive Detection - Generally available for new critical and high-severity SAST findings. - Uses LLM-based agentic reasoning to assess whether a vulnerability is likely real or a false positive. - Adds: - A confidence score - An AI-generated explanation - A “Likely false positive” or “Likely real” badge - Findings can be filtered in the Vulnerability Report so teams can prioritize likely real vulnerabilities. - The assessment remains a recommendation that teams can review and audit. ## Agentic SAST Vulnerability Resolution - Currently in beta. - For findings judged unlikely to be false positives, the agent: - Reads the vulnerable code and its surrounding context - Generates a proposed fix - Validates the fix with automated tests - Opens a merge request for developer review - The merge request includes code changes, a confidence score, and an explanation of the remediation. - Developers should still carefully review AI-generated changes before merging. ## Secret False Positive Detection - Currently in beta. - Identifies likely test credentials, placeholder values, example tokens, and other dummy secrets. - Provides confidence scores, explanations, and visual badges in the Vulnerability Report. - Runs automatically on the default branch and can also be triggered manually with “Check for false positive.” - The goal is to help teams focus on exposed credentials that represent genuine risk. GitLab 18.10 extends AI assistance across the vulnerability lifecycle: filtering out misleading findings, explaining security assessments, and proposing validated code fixes. Teams using GitLab Ultimate and Duo Agent Platform can use these features to reduce review effort while retaining human control over security decisions and merges.

Read original(opens in new tab)
gitlab3 min readCurated summary

Code Review without the bottlenecks or the bill

GitLab presents Code Review Flow as an agentic review capability designed to reduce code-review bottlenecks caused by AI-accelerated development. It runs automated, context-aware reviews for every merge request at a flat cost of $0.25 each, enabling teams to review changes in parallel rather than rationing reviews or waiting hours in a queue. The post concludes that predictable pricing and organization-wide automation can make continuous AI review practical. ## The Code-Review Bottleneck - AI coding tools have increased development speed, but review capacity has not kept pace. - Code-review times reportedly rose 91% on teams using AI coding tools. - Engineers at large companies wait a median of 13 hours for pull requests to merge. - 44% of engineering teams identify slow reviews as their biggest delivery blocker. - Competing AI review tools may cost $15–$25 per review, encouraging teams to limit their use. ## How Code Review Flow Works - Automatically starts when a merge request is opened. - Scans the code changes and explores relevant repository context. - Checks pipeline status, security findings, and compliance requirements. - Produces structured inline feedback grounded in the project’s broader context, not only the diff. - Runs within GitLab and can process hundreds of reviews in parallel across an organization. - Supports project-specific review instructions and different agents, including Claude Code, Codex, or custom agents. ## Flat-Rate Pricing and Savings - Each review costs 0.25 GitLab Credits, equivalent to $0.25 at list pricing. - Four reviews cost one credit, regardless of merge-request size or complexity. - The predictable pricing avoids token estimates and unexpected bills. - Compared with an estimated $25 cost for 15 minutes of senior-engineer review time, the post claims a 99% reduction in per-review cost. - Parallel execution can unblock merge requests in minutes instead of hours. ## Enabling Reviews by Default - The low fixed price is intended to remove the need to prioritize which merge requests receive AI review. - Teams can enable Code Review Flow automatically for every merge request and project. - Automated reviews handle routine feedback while engineers focus on architecture, mentorship, and higher-level decisions. - Consistent project-specific guardrails can be applied across different review workflows. ## Availability - The pricing is available through GitLab Duo Agent Platform on GitLab.com and GitLab Dedicated. - Self-managed GitLab instances must run version 18.8.4 or later. - GitLab recommends starting a trial or contacting an account representative to enable the feature. Teams seeking to reduce review delays can consider enabling Code Review Flow by default, while still treating automated feedback as a supplement to human review for architectural and judgment-intensive changes.

Read original(opens in new tab)
gitlab3 min readCurated summary

GitLab 18.10: Agentic AI now open to even more teams on GitLab

GitLab 18.10 makes agentic AI available to Free GitLab.com teams without requiring a subscription upgrade. By purchasing shared monthly GitLab Credits, teams gain access to planning, code generation, automated code review, and pipeline troubleshooting. The update also introduces predictable flat-rate pricing for code reviews, while Premium remains attractive for teams needing broader platform capabilities and included credits. ## Agentic AI for Free-tier teams - Free top-level GitLab.com groups can purchase a monthly commitment of GitLab Credits through group billing. - Credits are shared across the entire team, so organizations pay for AI usage rather than per-user access. - Teams receive access to capabilities previously available to Premium and Ultimate customers, including: - Planner Agent - Developer Flow - Code Review Flow - Fix CI/CD Pipeline Flow - Agentic Chat - Code Suggestions - Custom agents and flows - Group owners can use the GitLab Credits dashboard to monitor which agents and workflows consume credits. ## From planning to deployment GitLab describes a workflow covering the full software lifecycle: - **Planner Agent** turns a natural-language feature request into structured issues with descriptions, labels, and relationships. - **Developer Flow** reads an issue, generates code, runs tests, and opens a merge request. - **Code Review Flow** performs multi-step automated reviews and posts inline feedback based on repository context and code changes. - **Fix CI/CD Pipeline Flow** analyzes failed job logs, identifies likely root causes, and proposes fixes. - Agentic Chat supports iterative tasks such as refactoring, extending, or explaining code. ## Flat-rate automated code review - Code Review Flow costs **0.25 GitLab Credits per review**, regardless of merge request size, repository complexity, or internal processing steps. - Four reviews consume one credit. - The fixed price makes costs easier to forecast for both small and high-volume teams. - Automated reviews can run concurrently, reducing review queues and freeing human reviewers to focus on architecture and business logic. ## Why Premium may be the next step - GitLab Premium costs **$29 per user per month** and includes 12 promotional credits per user. - A 20-person team would receive 240 credits monthly—enough for approximately 960 automated code reviews or a mixture of AI workflows. - Premium also adds advanced CI/CD, merge approvals, code owners, governance features, and unified project context. - Teams that begin with Free plus purchased credits may find Premium more economical as AI becomes central to their development process. ## Getting started Free GitLab.com teams can purchase credits through group billing and begin using the Duo Agent Platform immediately. Teams seeking broader collaboration, governance, and CI/CD features can instead evaluate GitLab Premium or Ultimate.

Read original(opens in new tab)
gitlab2 min readCurated summary

Agentic code reviews for $0.25 each

GitLab introduces Code Review Flow, an agentic AI review feature priced at a flat $0.25 per merge request. It automatically analyzes code, repository context, pipelines, security findings, and compliance requirements, producing structured inline feedback. The post argues that predictable pricing and parallel execution can reduce review costs and shorten merge queues. ## The Code Review Bottleneck - AI coding tools have increased development speed, but review capacity has not kept pace. - Code review times have reportedly risen 91% on teams using AI coding tools. - Engineers at large companies wait a median of 13 hours for pull requests to merge. - 44% of engineering teams identify slow reviews as their biggest delivery blocker. - Existing AI review tools often use unpredictable token-based pricing, with some costing $15–$25 per review. ## How Code Review Flow Works - It starts automatically when a merge request is opened. - The agent: - Scans the code changes. - Explores relevant repository context. - Checks pipeline status and results. - Reviews security findings and compliance requirements. - Produces structured inline comments. - Because it runs within GitLab, reviews can execute in parallel across projects and organizations rather than sequentially in individual developers’ environments. ## Flat-Rate Pricing and Savings - Each review costs 0.25 GitLab Credits, or $0.25 at list pricing. - The price is the same regardless of merge request size or complexity. - Four reviews cost one GitLab Credit, making usage easy to forecast. - Compared with an estimated $25 cost for 15 minutes of senior-engineer review time, GitLab claims a 99% reduction in per-review cost. - Parallel reviews can unblock merge requests within minutes instead of hours. ## Scaling Reviews Across Teams - The low fixed price makes it practical to run reviews on every merge request. - Teams can define project-specific review instructions and guardrails. - Different projects can use Code Review Flow, Claude Code, Codex, or custom agents. - Results remain visible in GitLab while reviews run concurrently. ## Availability - The $0.25 pricing is available on GitLab.com, Dedicated, and self-managed GitLab instances running version 18.8.4 or later. - Users can try GitLab Duo Agent Platform through a free trial or contact their GitLab account representative. The post recommends enabling automated reviews broadly rather than reserving them for high-priority changes, using AI to handle routine feedback while engineers focus on architecture, mentorship, and higher-value decisions.

Read original(opens in new tab)
meta3 min readCurated summary

Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels surface videos that friends have liked or interacted with, combining content discovery with opportunities for conversation. The system uses machine-learning models to estimate viewer-friend closeness, retrieve relevant friend-interacted videos, and rank them alongside conventional recommendation signals. Its goal is not to show the most bubbles possible, but to identify meaningful connections and content that can drive both engagement and social interaction. ## System Architecture - The recommendation system combines: - **Viewer-friend closeness**, determining whose interactions matter most. - **Video relevance**, determining which friend-interacted videos best fit the viewer. - Multiple friends interacting with the same video can indicate stronger shared interest. - Social discovery and engagement reinforce one another: relevant friend content encourages interaction, which improves the system’s understanding of the social graph. ## Modeling Viewer-Friend Closeness - Facebook uses two complementary models: - A survey-based model estimating real-world relationship strength. - An activity-based model estimating closeness from on-platform behavior. - The survey model considers: - Mutual friends and interaction patterns. - User-provided attributes such as location. - Number of friends and posts shared. - Communication frequency and other survey proxies for offline closeness. - Users are periodically asked whether they feel close to a randomly selected connection. - The model is refreshed regularly and performs weekly inference across trillions of friend relationships. - The activity-based model learns from likes, comments, reshares, and interactions occurring after bubbles are shown. - Facebook prioritizes connection quality over quantity: larger friend networks may create more opportunities, but the system aims to surface only relationships likely to make recommendations meaningful. ## Retrieving and Ranking Friend Content ### Expanding Candidate Retrieval - The retrieval stage explicitly sources videos interacted with by close friends. - This expands the recommendation funnel so high-quality friend content can reach downstream ranking systems. - Without dedicated retrieval, relevant friend videos might never become candidates. ### Adding Social Context to Ranking Models - Friend-interacted videos could rank poorly when models lacked viewer-friend closeness information. - The system added bubble interaction signals and relationship-strength features to early- and late-stage multi-task, multi-label ranking models. - These features help models distinguish social relevance from ordinary content-interest signals. - Feedback from bubble impressions and resulting interactions continuously flows back into model training. - Ranking objectives consider: - Watch time. - Likes and comments. - The probability of engagement after a bubble impression: `P(video engagement | bubble impression)`. - Tunable weights balance entertainment and video quality against social goals such as discovering friends’ interests and encouraging conversation. ## Client Infrastructure and Reels Performance - Friend-bubble metadata had to be integrated without harming Reels’ core experience. - The implementation targeted: - Smooth scrolling. - No additional loading latency. - Low CPU usage during metadata retrieval and processing. - Facebook aligned bubble metadata retrieval with the existing video prefetch window, which already loads metadata, thumbnails, and buffered content before playback. - This allows the system to reuse cached results and avoid adding unnecessary work during scrolling. Friend bubbles work best when social relevance and content quality are optimized together. By combining relationship models, friend-aware retrieval and ranking, feedback-driven learning, and performance-conscious client infrastructure, Facebook turns shared video interests into lightweight opportunities for discovery and conversation.

Read original(opens in new tab)
aws2 min readCurated summary

Our First 2026 Heroes Cohort Is Here! | Amazon Web Services

AWS has announced its first 2026 Heroes cohort, recognizing Maurizio, Ray Goh, and Sheyla Leacock for combining technical expertise with community leadership. Their work spans cloud architecture, generative AI, machine learning, and cybersecurity, while emphasizing mentorship, education, and meaningful human connections. Together, they demonstrate how technology leaders can expand access to skills and strengthen communities globally. ## Maurizio – Pignola, Italy - CTO and organizer of the AWS User Group Basilicata. - Has spent more than a decade developing cloud communities and technology ecosystems in areas where they previously did not exist. - Founded an international technology conference in a small mountain village, connecting global experts with local developers. - Covers topics including cloud architecture, DevOps, and web scaling, alongside creative networking opportunities. - Mentors children, university students, and professionals transitioning into cloud careers. - Combines technical leadership with inclusive, cross-generational community building. ## Ray Goh – Singapore - AI and machine learning community leader involved in AWS programs since 2018. - Founded The Gen-C in 2024, offering public library workshops on generative AI, LLM fine-tuning, and AWS AI agents. - Has spoken at major AWS events and contributed to the AWS Machine Learning Blog. - Led DBS Bank’s AWS DeepRacer initiative, which trained more than 3,100 employees. - Trained over 1,300 ASEAN students in LLM techniques in 2025. - Supports skills-based programs teaching AI and machine learning to women, children, and young people. ## Sheyla Leacock – Panama City, Panama - IT security professional, mentor, technical writer, and international speaker. - Leads the AWS User Group in Panama and participates in AWS Community Days and regional meetups. - Has spoken at AWS Summits, AWS re:Invent PeerTalk sessions, and more than 20 international conferences. - Publishes educational content focused on AWS cloud computing and cybersecurity. - Works with universities as a guest lecturer to help develop future technology and security professionals. - Strengthens the cloud and cybersecurity ecosystem through education, knowledge sharing, and community leadership. The new cohort highlights the broader impact of community-driven technology leadership. Readers can visit the AWS Heroes webpage to learn more about the program or connect with a Hero.

Read original(opens in new tab)
cloudflare2 min readCurated summary

Introducing Custom Regions for precision data control

Cloudflare is expanding Regional Services with three new managed regions—Turkey, the UAE, and IRAP—and introducing Custom Regions. The feature lets customers define exactly where TLS termination and Layer 7 processing occur while still using Cloudflare’s global network for traffic ingestion and DDoS protection. This combines local data-sovereignty control with globally distributed security and performance. ## Global Security with Local Compliance - Traffic enters through the nearest Cloudflare data center, where Cloudflare applies global-scale Layer 3 and Layer 4 DDoS protection. - Before decryption, request metadata is inspected and traffic is routed over Cloudflare’s private backbone to a data center inside the customer’s designated region. - TLS termination and Layer 7 services—including WAF, Bot Management, and Workers—run only within that region. - After processing, traffic is re-encrypted and sent securely to the origin. - This design localizes data inspection without forcing customers to sacrifice Cloudflare’s global attack-mitigation capacity. ## Expanded Cloudflare Managed Regions - Regional Services originally supported the EU, UK, and U.S. - Cloudflare now offers 35 predefined regions. - Newly added options include: - Turkey - United Arab Emirates - IRAP, supporting Australian compliance - ISMAP, supporting Japanese compliance ## Custom Regions - Customers can define their own geographical boundaries instead of selecting only Cloudflare-managed regions. - Custom Regions support: - Individual countries - Arbitrary combinations of countries - Regions that exclude specified countries - Early-access use cases include: - Keeping AI prompts and responses within selected countries - Running country-specific promotions - Meeting government contractual requirements - Aligning regions with corporate structures such as EMEA, MENA, or APAC - Example definitions include North America, everywhere except North America, or countries where Fahrenheit is commonly used. ## How Region Enforcement Works - A region represents a set of Cloudflare data centers. - Managed regions use Cloudflare-defined membership. - Custom Regions use expressions, commonly based on the data center’s ISO country code: - `country_code == "TR"` selects Turkey. - `country_code in ["DE", "FR", "NL"]` selects Germany, France, and the Netherlands. - Negated expressions can exclude specific countries. - The same Regional Services architecture applies regardless of who defines the boundary; Custom Regions simply give the customer control over the membership rules. Custom Regions are best suited to organizations with precise sovereignty, compliance, performance, or organizational requirements. They provide fine-grained control over where sensitive traffic is decrypted and processed while preserving Cloudflare’s global network protections.

Read original(opens in new tab)
line4 min readCurated summary

Large-scale iOS Settings System Unraveled through AttributedString Structure

LINE’s “Service Configuration” system lets teams deploy features dynamically without waiting for LINE’s two-week app release cycle. As the iOS app grew to roughly 700 configuration keys across 60 modules, its monolithic design created dependency, usability, concurrency, testing, and QA problems. The article argues that the original design was reasonable at small scale but needed to evolve, beginning with lessons from Foundation’s type-safe `AttributedString` design. ## What Service Configuration Provides - Service operators modify values through an administration page. - The server notifies LINE clients, which fetch updated values. - Values are selected based on factors such as: - User region - Device - OS version - The system supports: - Feature flags - Rollbacks - A/B tests - Error-reporting sample rates - UI behavior policies - Configuration is delivered as a string-to-string dictionary, for example: - `"function.media.image_medium": "1280,70"` - `"function.media.message.flow.v2.image": "Y"` ## Problems Caused by the Monolithic Design The original implementation required every key to be declared in one roughly 7,000-line file. Although this was simple initially, growth in teams and modules made the structure increasingly costly. ### Circular Dependencies and Weak Typing - Configuration values were exposed as raw strings because the configuration module could not depend on feature-specific modules. - For example, `"1280,70"` represented image dimensions and JPEG quality, but callers had to parse it into an `ImageTransferQuality` value themselves. - Defining `ImageTransferQuality` in the configuration module avoided repeated parsing but polluted unrelated modules with photo-specific types. - Defining it in the photo module preserved separation of concerns but created an impossible reverse dependency. ### Incomplete and Confusing Abstractions - Developers had to understand server-specific encoding rules and implementation details. - Boolean values were sent as `"Y"` and `"N"`, requiring a custom `decodeBoolIfPresent(forKey:)` method. - The custom decoder’s name resembled Swift’s standard decoding API, making incorrect implementations easy to write and review. - Decoding failures could silently fall back to defaults, making the underlying problem difficult to diagnose. - The same default value often had to be declared three times: - A property-group default - A decoding fallback - A global `defaultConfiguration` entry - These duplicated defaults served subtly different purposes, although the distinctions were generally unnecessary. ### Lack of Thread Safety - Configuration groups were lazily decoded and replaced when new server values arrived. - Multiple services could read configuration values concurrently on different threads. - This caused use-after-free crashes— reportedly hundreds per day—leading to bug tickets and hotfix releases. - As the number of services and concurrent operations increased, this became a systemic issue rather than an occasional edge case. ### No Built-in Debug Overrides - QA frequently needed to temporarily change configuration values. - Because the system had no override mechanism, each feature required custom: - Persistent storage - Debug-menu UI - Value-display text - Implementing this repeatedly required edits across several files and modules. ### Fragmented Test Doubles - Since `LineConfigurationManager` was a singleton, modules created narrow protocols and custom mocks for the settings they used. - This resulted in dozens of duplicated protocols and test doubles. - These had to be updated alongside configuration keys and could fall out of sync. - Differences between mocks and production behavior could allow bugs to escape tests or create false failures. ## Looking to Established Designs The team first distilled the required properties of a replacement: - Type-safe access to a large number of key-value pairs - Independent key definitions by each module - Safe behavior under concurrency They identified Foundation’s `AttributedString` as a useful precedent because it manages many typed attributes while allowing UIKit, AppKit, SwiftUI, and other frameworks to define their own attributes independently. The article presents this as the starting point for redesigning Service Configuration around a more modular and type-safe architecture.

Read original(opens in new tab)
stripe2 min readCurated summary

Introducing the Machine Payments Protocol

AI agents are moving beyond chatbots toward autonomous systems that plan, act, and evaluate results, creating demand for agent-friendly commerce. Stripe and Tempo are launching the Machine Payments Protocol (MPP), an open standard that lets agents make programmatic payments to businesses and services. MPP supports microtransactions, recurring payments, stablecoins, fiat, and existing Stripe payment methods without requiring human intervention. ## The Challenge of Agent Payments - Traditional financial workflows are designed for humans. - Agents often cannot independently: - Create accounts - Navigate pricing and subscription options - Enter payment details - Configure billing - These obstacles limit agents’ ability to purchase services and participate in the internet economy. ## How the Machine Payments Protocol Works - An agent requests a resource from a service, API, MCP server, or other HTTP endpoint. - The service returns a payment request. - The agent authorizes payment. - The requested resource is delivered automatically. - Stripe businesses can integrate MPP through the PaymentIntents API with only a few lines of code. - Payments appear in Stripe’s existing API and Dashboard and settle through the business’s normal balance, currency, and payout schedule. - Standard Stripe capabilities remain available, including tax calculation, fraud protection, reporting, accounting integrations, and refunds. ## New Agentic Business Models MPP is already enabling agents to pay for services such as: - Browserbase: headless browsers billed per session - PostalForm: printing and mailing physical documents - Prospect Butcher Co.: ordering food for pickup or delivery in New York City - Stripe Climate: making programmatic contributions - Parallel Web Systems: paying per API call for web access Payments can use stablecoins, cards, buy now, pay later methods, and Shared Payment Tokens. ## Stripe’s Agent Economy Infrastructure Stripe positions MPP alongside its broader Agentic Commerce Suite, Agentic Commerce Protocol, MCP integrations, and support for x402. Together, these tools are intended to help businesses sell directly to agents and support new automated commerce patterns. Businesses interested in enabling agent payments can review Stripe’s MPP documentation and join the early-access program.

Read original(opens in new tab)