Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

The truly programmable SASE platform

Cloudflare argues that true SASE programmability goes beyond APIs, Terraform, webhooks, and alerts. It means intercepting security events, enriching them with external context, and making real-time decisions through custom logic. By running Cloudflare One and its Developer Platform on the same global edge network, Cloudflare aims to let customers apply programmable, low-latency policies without stitching together separate infrastructure. ## What “Programmability” Means - Traditional programmability supports configuration and automation, such as sending Slack alerts when policies trigger. - True programmability allows security systems to: - Inspect an event before access is granted. - Query external systems for additional context. - Make or change an access decision in real time. - Example: a request to a regulated application could be checked against a learning management system to confirm that the user’s compliance training is current. Expired or missing certification would result in denial and redirection to training. ## Cloudflare’s Programmable SASE Architecture - Cloudflare’s network spans more than 330 cities and reaches approximately 95% of Internet-connected users within 50 milliseconds. - Cloudflare One and the Developer Platform run on the same infrastructure and use shared network primitives. - This enables Workers to extend inline services such as Access without requiring separate cloud infrastructure. - Customers can: - Call external risk APIs. - Add dynamic request headers. - Validate browser attributes. - Route traffic according to custom business logic. - Running custom logic at the edge reduces latency and avoids the operational overhead of webhook-based integrations and disconnected systems. ## Custom Actions in Security Policies - Conventional gateways generally limit policy outcomes to actions such as allow, block, isolate, or quarantine. - Cloudflare is expanding policies to support managed and custom actions. - Potential uses include: - Injecting headers based on user identity claims. - Obtaining real-time verdicts from external risk engines. - Restricting access based on location or working hours. - Updating risk lists based on scheduled analysis of user activity. - Custom actions can invoke a Worker when a Gateway HTTP policy matches, giving the code access to request context and allowing decisions in milliseconds. - Managed actions will offer templates for common use cases such as IT service management, redirects, and compliance workflows. ## Automated Device Session Revocation - One customer needed periodic re-authentication for Cloudflare One Client users, similar to traditional VPN session expiration. - Cloudflare’s built-in session controls were application-specific rather than global and time-based. - The customer implemented a scheduled Worker that: - Queries the Cloudflare Devices API. - Handles cursor-based pagination to retrieve registrations. - Calculates how long each device has been inactive. - Deletes registrations exceeding a configured inactivity threshold. - Forces affected users to authenticate again through their identity provider. - The example also supports environment-based configuration and a dry-run mode for testing before revocations are applied. Cloudflare’s recommendation is to treat SASE policies as programmable decision points rather than fixed allow-or-block rules. Combining Cloudflare One with edge Workers can provide faster, more context-aware security automation while reducing integration complexity.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Modernizing with agile SASE: a Cloudflare One blog takeover

The post argues that changing work patterns, AI agents, and Internet-based perimeters require organizations to move beyond fragmented legacy networks toward “agile SASE.” It presents Cloudflare One as a composable platform that combines networking and security on a global connectivity cloud, using single-pass processing to avoid service-chaining bottlenecks. Cloudflare positions this architecture as a faster path to modernization, beginning with focused use cases rather than a large-scale “big bang” migration. ## The Case for Agile SASE - Hybrid work and AI-driven traffic are making traditional corporate perimeters and office-based networks obsolete. - Legacy firewalls, VPN concentrators, and hardware appliances create a “fragmentation penalty.” - Technical debt accumulates through: - Conflicting firewall rules - Manual patching - Aging hardware - Infrastructure unable to handle AI-scale traffic - First-generation SASE platforms often shifted fragmentation from physical hardware into isolated cloud and operational silos. - The resulting problem is not a shortage of security data, but difficulty enforcing consistent policies across a borderless enterprise. ## Cloudflare One’s Single-Pass Architecture - Cloudflare describes zero trust as the security model and Cloudflare One as the platform for implementing it. - The platform converges networking and security through a global connectivity cloud spanning more than 300 cities. - Security checks can run simultaneously on every server rather than processing traffic sequentially through separate services. - This avoids service chaining, which can introduce latency and operational complexity. - Cloudflare frames the result as a programmable, composable platform rather than a collection of acquired or loosely connected tools. ## Five Themes for Network Modernization The company’s planned technical series focuses on: - **A new network standard:** Building a programmable, future-ready Internet foundation. - **Identity beyond passwords:** Combining human and device verification instead of relying only on credentials. - **Signal over noise:** Using AI to convert large volumes of security data into actionable guidance. - **The autonomous edge:** Improving performance and reducing friction as part of the security strategy. - **A unified vision:** Showing how enterprises and partners can standardize on Cloudflare One at scale. ## Programmability and Developer Integration - Cloudflare One runs alongside Cloudflare Workers, allowing teams to write code that responds to security events in real time. - This extends policy enforcement beyond simple allow/block decisions. - Organizations can automate more sophisticated security and operational workflows. - Cloudflare argues that this flexibility helps technology teams support faster business growth while improving protection. ## Practical Starting Points The post recommends beginning with focused modernization projects: - **Remote access:** Replace maintenance-heavy VPNs with clientless access and faster zero trust adoption. - **Email protection:** Detect phishing, Business Email Compromise, and related multi-channel threats with AI-powered controls. - **DNS filtering:** Block malicious websites for hybrid workers using DNS protection based on the 1.1.1.1 resolver. - **AI governance:** Identify shadow AI usage and control how organizational data enters generative and agentic AI systems. - **Branch networking:** Treat offices and coffee-shop workspaces as remote sites, reducing dependence on dedicated hardware. Organizations evaluating SASE should favor platforms that provide consistent policy enforcement, programmable controls, and incremental adoption. Cloudflare’s recommended entry point is to start with a specific need—such as VPN replacement or AI governance—and expand toward a unified connectivity and security architecture.

Read original(opens in new tab)
netflix3 min readCurated summary

Mount Mayhem at Netflix: Scaling Containers on Modern CPUs

Netflix’s effort to modernize its container runtime exposed a hardware-level bottleneck rather than an application problem. Under heavy startup concurrency, containers with many image layers triggered massive mount and unmount activity, causing kernel lock contention, systemd stalls, and container startup failures. The issue was especially severe on older dual-socket NUMA instances, while newer single-socket systems scaled much more reliably. ## Container Startup at Netflix - New AWS capacity is rapidly filled with pods as applications scale. - Some nodes became unresponsive, with: - Health checks timing out for more than 30 seconds - Kubelet requests to containerd timing out - systemd processing huge numbers of mount events - The mount table taking tens of seconds to read - The problem primarily affected `r5.metal` instances running images with more than 50 layers. ## Mount Lock Contention - With user namespaces, containerd performs several mount operations for every image layer: - `open_tree()` references the layer. - `mount_setattr()` applies the container’s ID mapping. - `move_mount()` creates an ID-mapped bind mount. - These bind mounts become OverlayFS lower directories and are later unmounted. - The Linux VFS uses global mount-related locks, so concurrent container creation causes CPUs to contend on the same kernel locks. - For 100 containers with 50 layers each, containerd performs the process twice: - `100 × 2 × (1 + 50 + 50) = 20,200` mount operations - This makes startup cost depend heavily on both container concurrency and image layer count. ## Why the New Runtime Exposed the Problem - The old Docker-based runtime shifted file ownership while unpacking images. - All containers shared one host user range, avoiding repeated per-container mount work. - The new containerd-based runtime assigns each container a unique host user range for stronger isolation. - Instead of rewriting file ownership during extraction, it uses Linux ID-mapped mounts to apply ownership mappings efficiently. - This improves security and avoids expensive image copying, but creates many additional mount operations during startup. ## Differences Between AWS Instance Types Netflix compared: - `r5.metal`: 5th-generation Intel, dual-socket, multiple NUMA domains - `m7i.metal-24xl`: 7th-generation Intel, single-socket, single NUMA domain - `m7a.24xlarge`: 7th-generation AMD, single-socket, single NUMA domain Results showed: - At low concurrency—around 20 containers or fewer—all systems performed similarly. - `r5.metal` began failing at roughly 100 concurrent container launches. - Newer Intel instances maintained lower startup times and better success rates. - AMD-based `m7a` instances scaled most consistently and had the fewest failures. ## Kernel and CPU-Level Diagnosis - Profiling showed that containerd spent most of its time in Linux VFS path lookup code. - Specifically, threads were spinning in `path_init()` while waiting on a sequence lock. - Intel Topdown Microarchitecture Analysis found: - 95.5% of pipeline slots stalled on contested accesses - 57% attributed to false sharing - Cache-line bouncing and global lock contention, rather than raw CPU capacity, dominated performance. ## NUMA as a Contributing Factor - NUMA systems divide memory among processor sockets. - Local memory access is faster, while remote access crosses an interconnect and introduces additional latency. - The dual-socket layout of `r5.metal` amplified contention around shared mount-related data. - The better behavior of newer single-socket instances indicated that CPU topology and memory locality were key contributors to the container startup bottleneck. ## Practical Conclusion High-concurrency container launches can overwhelm kernel mount infrastructure, especially when using per-container ID mapping and images with many layers. Netflix’s results suggest minimizing image layers, controlling startup concurrency, and favoring newer single-socket hardware can substantially improve reliability and scaling.

Read original(opens in new tab)
microsoft3 min readCurated summary

Engineering and algorithmic interventions for multimodal post-training at Microsoft scale

At production scale, post-training multimodal agents fail for reasons that standard reinforcement-learning literature often overlooks. Heterogeneous tasks, long tool-use trajectories, noisy reward sources, and strict latency and safety requirements can make aggregate reward look healthy while the policy gradient becomes uninformative and important capabilities regress. The post presents interventions designed to preserve useful advantage signals as scale, task diversity, and interaction horizons grow. ## Production-Scale Challenges - Copilot agents must simultaneously handle: - Tool orchestration - Enterprise documents and mixed-media inputs - Content moderation - Multi-step execution - Trajectories range from roughly 100 to more than 2,000 tokens and span 6 to 25+ interaction steps. - Rewards come from programmatic checks, human judgments, and implicit usage signals, each with different noise and latency. - A single scalar reward can hide regressions in robustness, long-horizon planning, or downstream task success. - Aggregate reward may rise while gradient updates increasingly depend on a small, unrepresentative subset of trajectories. ## Staged Objective Curriculum - The team separates: - **Verifiable objectives**, such as tool syntax and format compliance - **Preference objectives**, such as tool choice and response quality - Training uses only verifiable objectives during the first 30%. - Preference signals are then introduced linearly. - An entropy floor, implemented through a KL penalty activated below a threshold, prevents premature policy collapse. - Entropy bonuses were insufficient because the issue was not simply exploration; optimization was favoring behaviors that were easy to score. - A 30% warmup worked better than 10% or 50% across task families. - Early text-only supervision could also activate multimodal capabilities more reliably than noisy direct multimodal supervision, assuming adequate cross-modal alignment from pretraining. ## Adaptive Curriculum Based on Estimator Health - The team monitors effective sample size (ESS): `ESS = (Σ wᵢ)² / Σ wᵢ²` - ESS measures how many trajectories meaningfully contribute after importance weighting. - ESS falling below 20% of nominal batch size predicted learning stalls by about 35 epochs. - When ESS drops, the system: - Injects near-miss trajectories from a reservoir buffer - Temporarily increases the KL penalty to limit policy drift - Near-misses worked better than hard negatives because they preserve useful distinctions near the decision boundary instead of merely pushing the policy away from failure. - The intervention maintained ESS above 70%, with approximately 15% additional memory usage. ## Variance-Corrected Normalization - Per-task gradient normalization balances task magnitudes but ignores variance within each task. - Broad categories such as “coding” may contain trajectories ranging from 100 to 2,000 tokens, with very different variance. - Importance weighting can cause long trajectories to dominate the effective gradient even after task-level normalization. - The excerpt ends while introducing the team’s variance-correction approach, so its implementation and results are not included here. The central recommendation is to treat estimator health—not just reward and task metrics—as a first-class training signal. Monitoring ESS, controlling objective timing, and accounting for trajectory variance can help prevent healthy-looking dashboards from masking policy collapse and capability regressions.

Read original(opens in new tab)
pinterest3 min readCurated summary

Bridging the Gap: Diagnosing Online–Offline Discrepancy in Pinterest’s L1 Conversion Models

Pinterest found that strong offline gains in L1 conversion-rate models did not translate into online improvements because training and serving environments were not aligned. Although experimental models reduced LogMAE by roughly 20–45% and improved calibration, online A/B tests showed neutral or worse CPA and unexpected oCPM mix shifts. The investigation identified feature coverage gaps and embedding version skew as structural causes rather than problems with offline evaluation or serving reliability. ## How L1 Models Are Evaluated - L1 filters and prioritizes ads under strict latency limits before downstream ranking and auction stages. - Offline evaluation focused on: - LogMAE and calibration - Performance across candidate pools and pCVR percentiles - Multiple data sources, including auction winners and candidates - Online evaluation focused on: - CPA and other business metrics - Candidate counts and recall across funnel stages - Differences among optimization types, especially oCPM traffic ## Hypotheses That Were Ruled Out - **Offline evaluation errors** - The experimental model consistently outperformed production across three log sources. - Gains remained across pCVR buckets, including after outlier handling. - **Exposure bias** - Increasing treatment traffic from approximately 20% to 70% did not resolve the online over-calibration issue. - **Serving failures** - Control and treatment had comparable success rates and p50/p90/p99 latency. - Timeouts and tail latency were therefore unlikely to explain the discrepancy. ## Missing Features in L1 Serving - Offline training used rich logged features, while online L1 embeddings only included features explicitly onboarded into the embedding pipeline. - Important feature families were absent online, including: - Targeting specification flags - Offsite conversion visit counts over 1-, 7-, 30-, and 90-day windows - Annotations and MediaSage image embeddings - Models learned to depend on these signals during training, but received a substantially thinner feature set when serving many oCPM and performance-oriented ads. - Pinterest updated UFR configurations to add the missing features to L1 embeddings. - Online feature coverage recovered, and online loss improved for CVR and engagement models, particularly on shopping traffic. - UFR tooling was also changed so features onboarded for L2 are automatically considered for L1 embedding usage. ## Query–Pin Embedding Version Skew - Pinterest’s two-tower architecture requires query and Pin embeddings to be generated from compatible model checkpoints. - Offline evaluation generally uses one fixed checkpoint for both towers. - Online pipelines could instead serve query and Pin embeddings produced from different model versions, creating a mismatch between training assumptions and production behavior. - This version skew was identified as a second structural source of online–offline inconsistency. ## Practical Conclusion Offline model quality is not sufficient for launching L1 improvements. Teams must verify feature coverage in serving artifacts such as ANN indices, enforce synchronized query and Pin embedding versions, and monitor funnel behavior and online feature coverage alongside standard offline metrics.

Read original(opens in new tab)
github3 min readCurated summary

From idea to pull request: A practical guide to building with GitHub Copilot CLI

GitHub Copilot CLI helps developers move from an idea to reviewable code without leaving the terminal. The recommended workflow is to begin with intent, let Copilot propose plans and scaffolding, validate changes through tests and diffs, then move to an IDE for refinement and GitHub for collaboration. Copilot accelerates development but does not replace design judgment, code review, or user approval. ## What Copilot CLI Is—and Isn’t - It is a GitHub-aware coding agent that operates in the terminal. - Developers can describe goals in natural language and use `/plan` or `Shift + Tab` planning mode. - It proposes commands, file changes, and diffs for review before execution. - It can generate files, modify code, and explain failures. - It does not silently run commands or eliminate the need for careful design and review. ## Start with Intent - Begin by describing the application or feature rather than choosing a framework or copying a template. - For example, ask Copilot to create a small web service with a JSON endpoint and tests. - Copilot may suggest a technology stack, file structure, and setup commands. - Review these suggestions before deciding what to execute. ## Scaffold Only What You Own - Once the direction is clear, ask Copilot to create a minimal project structure. - It can generate directories, configuration, test runners, and README files. - Generated scaffolding should be treated as a starting point, not an unquestioned design. - Developers remain responsible for reviewing, editing, or discarding the result. ## Iterate from Real Failures - Run tests directly within the CLI and use the resulting output as context. - Ask Copilot to explain a failure or propose a fix with a reviewable diff. - The recommended loop is: run a command, inspect the output, ask for help, and review the proposed change. - Use `explain` when understanding is the goal and `suggest` when seeking a concrete proposal. ## Handle Mechanical Repository-Wide Changes - Copilot CLI is effective for clearly scoped, repetitive work such as renaming symbols across a repository. - It can update related tests and provide a concrete diff. - Mechanical changes are relatively easy to inspect, revert, and validate. ## Move to the IDE for Precision - The terminal is best for fast exploration, planning, scaffolding, and low-ceremony changes. - Move to an editor or IDE when refining APIs, handling edge cases, and making design decisions. - A practical division is: - **CLI:** plan, generate diffs, and move quickly. - **IDE:** refine logic and shape the implementation. - **GitHub:** commit, open pull requests, review, and collaborate. ## Finish by Shipping on GitHub - Copilot CLI can help add descriptive commits, push changes, and create pull requests. - Pull requests make the work durable through teammate review, CI testing, and asynchronous collaboration. - The workflow can also add Copilot as a reviewer. - The ultimate value comes from reaching commits and pull requests, not merely generating suggestions. Copilot CLI is most effective as a momentum tool: use it to turn intent into concrete, testable changes, while retaining human control over design, approval, and review.

Read original(opens in new tab)
daangn5 min readCurated summary

Things I learned using 2

Karrot’s Taxonomy team built an LLM-powered system to classify marketplace posts, group activities, and local businesses into a shared category and attribute structure. After finding that manually managed taxonomies and event-only pipelines were difficult to scale, they created a configurable Taxonomy Management System using Dataflow/Beam, BigQuery, Kafka, and multiple LLM strategies. The system emphasizes scalable inference, rapid evaluation, multilingual support, and continuous taxonomy expansion. ## What a Taxonomy Is and Why It Matters - A taxonomy is a hierarchical category system, such as `Outerwear > Padding/Down > Long Padding`. - It can also include attributes that describe an item’s characteristics: - Category: long padding - Attributes: brand=Nike, color=black, material=polyester - A consistent taxonomy acts as a shared language across: - Search, including parent and child-category expansion - Recommendations and diversity controls - Advertising and targeting segments - Analytics and machine-learning features ## Karrot’s Taxonomy Challenges - Karrot manages roughly 1,400 marketplace categories across up to three levels. - Users are not required to manually select highly detailed categories because that would increase posting friction and produce unreliable labels. - Earlier systems used a Golang Kafka consumer to receive posting events and extract categories with an LLM. - This approach had several limitations: - Taxonomy definitions were managed separately by different teams. - Categories alone could not express useful properties such as season or material. - Batch processing and backfilling were difficult. - Expanding to data sources outside Kafka was inconvenient. - Quality monitoring and failure handling were insufficient. - Changes to prompts or models required slow offline and online experiments. ## The Taxonomy Management System - The new system centrally manages taxonomies, performs LLM-based classification, delivers category and attribute results, and monitors quality. - Dataflow with Apache Beam was selected because it supports: - Parallel, high-throughput LLM inference - Both streaming and large-scale batch processing - Existing team expertise compared with alternatives such as Spark or Flink - BigQuery serves as the source of truth for inference results. - Analysts and data scientists can query results directly. - Online consumers can receive results through Kafka sinks into the internal feature platform. ## Configuration-Driven and Extensible Design - Taxonomy definitions are stored in YAML, allowing different services and category trees to use the same framework. - Pipeline settings, worker sizing, Kafka topics, and BigQuery destinations are also configured through YAML. - LLM models and inference strategies can be selected through configuration, including: - Primary and evaluation models - Single-shot or two-stage categorization - Attribute extraction modes - Evaluation sampling ratios - The system is designed for multilingual taxonomies. - Large translation jobs are divided into chunks. - One LLM generates translations and another validates consistency and naturalness. - A depth-first traversal carries parent-category translations into child-category prompts to maintain terminology consistency. ## Creating and Expanding Taxonomies with LLMs - New taxonomies are developed by researching established taxonomies and generating candidate trees from real data. - Existing taxonomies are expanded by: - Classifying sampled data against the current taxonomy - Asking the LLM to suggest categories for unsuitable examples - Merging similar suggestions using LLM similarity judgments - Promoting sufficiently strong candidates for review - Candidates undergo two evaluations: - Whether the originating examples are correctly assigned to the new category - Regression testing comparing classifications under the old and new taxonomies - This process enabled the team to move beyond the existing 1,400 three-level categories and create taxonomies with more than 10,000 categories and six or more levels. ## LLM Categorization Strategies The team supports multiple strategies because the best approach depends on the model and taxonomy size: - **Single shot:** Provide all categories and ask the model to choose one. - **Hierarchical classification:** Select the best category at each depth, then continue through the chosen branch. - **Two-stage tournament:** Split categories into chunks, select candidates from each chunk, and run a second selection among those candidates. - Categorization and attribute assignment are separate modular Beam `DoFn` stages: - `Article → Category inference → Attribute inference` - New approaches can be added as interchangeable strategies without redesigning the whole pipeline. ## Evaluation with LLM-as-a-Judge - A sample of production data is processed by multiple different models. - Their labels are combined into a ground-truth label, generally through majority voting. - Each model’s output is compared against that ground truth. - Accuracy changes are tracked whenever the team modifies: - The LLM model - Prompts - Pipeline structure - Categorization or attribute strategies - The ground-truth method varies depending on whether the task involves: - A single category - Multiple categories - Multi-label attributes - Category quality is measured as a precision-at-one-style accuracy: the primary model’s category must match the ground-truth category. - Attributes are evaluated with precision and recall because a post can legitimately contain multiple attribute-value pairs. The main recommendation is to treat LLM classification as a production data pipeline rather than a one-off prompt: centralize taxonomy management, support both batch and streaming execution, make inference strategies configurable, and build automated evaluation and monitoring into the system from the beginning.

Read original(opens in new tab)
daangn5 min readCurated summary

How will long-term user modeling

Long-term user modeling captures persistent interests, cross-vertical behavior, and signals beyond what short-term recommendation logs can reveal. 당근 built a Transformer-based user encoder that learns from tens of billions of actions across its local marketplace, jobs, real estate, and other services, then exposes the resulting embedding as a shared feature for ranking, retrieval, and advertising models. The approach improved scalability and reuse, but introduced freshness and representation-transfer limitations. ## Why Long-Term User Modeling Matters - Recent actions reveal immediate intent, but miss recurring interests such as seasonal shopping or repeated moving-related searches. - Long-term, cross-vertical activity can connect behaviors such as: - Searching for real estate - Looking for furniture and appliances - Reading neighborhood moving advice - Longer histories can reduce selection bias caused by training only on items previously exposed by recommendation models. - Simply adding more history is insufficient because ranking systems are latency-sensitive and long sequences increase computation and infrastructure complexity. ## Shared User Embeddings as a Common Feature - A separate user encoder processes long-term history offline. - Home-feed ranking, candidate generation, and advertising models consume the resulting embedding as a shared user feature. - Benefits: - Downstream models avoid directly processing massive histories. - The encoder can scale independently in model size, data, and compute. - One embedding can be reused across multiple recommendation surfaces. - Limitations: - A downstream model receives only a fixed vector, so it cannot fully exploit the encoder’s richer representations. - Batch inference means recent actions are not reflected immediately. - Possible future improvements include more frequent or real-time updates, fine-tuning, and distillation. ## Contrastive User Modeling - The encoder uses a two-tower architecture: - A causal Transformer converts the user’s action sequence into a user embedding. - An MLP converts item features into item embeddings. - InfoNCE loss trains the user embedding to predict the next interacted item. - In-batch negatives provide alternative items for contrastive learning. - Training uses clicks and conversion actions across all major verticals and surfaces. - The dataset contains tens of billions of actions—around 150 times more than the existing home-feed candidate model’s training data. ## Item ID Embeddings vs. Content Embeddings ### Problems with Item ID Embeddings - New items have no learned ID embedding, creating a cold-item problem. - Hundreds of millions of item IDs require enormous embedding tables. - In the ID-based model, embedding tables accounted for over 99% of parameters, leaving little GPU capacity for the Transformer. - Hashing and embedding-sharding techniques were considered but did not provide a sufficient solution. ### Content Embeddings - The system switched to LLM-generated embeddings based on post metadata. - This enables: - Representations for newly created items - Much larger Transformer models, with Transformer parameters becoming roughly 1,000 times larger than in the ID-based setup - Large-scale lookup required two memory-efficient techniques: - `memmap` loads only needed embedding segments from disk and benefits from shared OS page caches during distributed training. - `bbhash` maps item IDs to embedding locations using roughly three bits per key, reducing mapping memory by about 97% compared with Python dictionaries. - Together, these methods made training with hundreds of gigabytes of item embeddings practical. ## Region-Constrained Batch Sampling - Standard in-batch negatives assume that other items in the batch were visible but not selected. - This assumption fails in a local service: users generally cannot view items outside their geographic area. - More than 86% of transactions occur within five kilometers, yet random batches mixed users and items nationwide. - Consequently, about 98% of random in-batch negatives were “impossible negatives”—items users could never have seen. - These negatives teach geographic unavailability rather than user preference, weakening the contrastive signal. ### RCBS Solution - Region-Constrained Batch Sampling (RCBS) groups users from the same region into a batch. - This reduced impossible negatives from 98% to 30%. - The remaining impossible negatives mainly came from differences in viewing radius or users’ historical activity in other regions. - Feasible negatives are harder because they represent items users could have viewed but rejected, forcing the model to distinguish genuine preferences among similar local items. ### Why Sampling Was Better Than Masking - Masking impossible negatives would remove most of the batch, drastically reducing effective batch size. - Hard-negative mining would require checking feasibility separately for each user and could be expensive and complex. - RCBS naturally produces more feasible and difficult negatives without changing the loss function or adding specialized mining. ## Applying the Embeddings - For home-feed and advertising ranking, the embedding is projected and concatenated with existing features. - The long-term encoder supplies persistent preference signals, while existing ranking models continue handling short-term and real-time signals. - For retrieval models, the best-performing approach used the user embedding alone to generate candidates rather than merely adding it as another feature. - The separate candidate source appeared to improve recommendation diversity. ## Embedding Refresh and Serving - Offline tests showed little difference between frozen embeddings and 12- or 24-hour refreshes. - Online A/B tests favored periodic updates, with shorter intervals performing better. - A 24-hour refresh cycle was selected as the best cost-performance trade-off. - GPU inference runs through a Beam pipeline on GCP Dataflow. - Only users who acted during the refresh window are reprocessed, avoiding unnecessary inference for inactive users. - Near-real-time inference remains a major future engineering challenge. The overall recommendation is to treat long-term user modeling as a separate, reusable representation system rather than forcing every downstream model to process extensive histories directly. For geographically constrained services, the training data pipeline—especially negative sampling—must reflect actual item visibility, making region-aware batching as important as the model architecture itself.

Read original(opens in new tab)
daangn4 min readCurated summary

Solving a 20

Carrot’s short-form video editor expanded its Android module to 200MB, later reduced to 40MB by moving filters and effects to a CDN. Because editing is used by relatively few users and is unavailable globally, the team adopted an on-demand Dynamic Feature Module (DFM) so most users would not install the extra assets. Rather than moving all editor code into the DFM, they ultimately separated only the large native libraries, achieving most of the size reduction with less operational risk. ## The Problem: A Large Video-Editing Module - Video editing introduced filters, effects, stickers, and native SDK resources totaling about 200MB. - CDN delivery reduced the module to 40MB and enabled updates and A/B testing without app releases. - The remaining 40MB was still an unnecessary burden for global users and users who never edit videos. ## Choosing On-demand Dynamic Feature Modules - Android App Bundles support: - **Install-time delivery:** installed with the app. - **On-demand delivery:** downloaded when requested. - **Conditional delivery:** installed based on conditions such as country or API level. - The team selected **On-demand Delivery** so the editor would download only when a user tapped the edit button. - After installation, the module remains local for subsequent use. - The user experience also required handling download progress, network failures, and first-use delays. ## Challenges of Separating the Entire Feature ### Hilt and Dependency Injection - Hilt combines dependency graphs at compile time, but the base and DFM modules are built separately. - DFM annotations such as `@Inject` and `@Module` are therefore unavailable during the base build. - The solution was to expose dependencies from the base module through a Hilt `EntryPoint`. - The DFM then constructs its own Dagger component at runtime. ### SplitCompat - On-demand modules are installed as split APKs. - The default `ClassLoader` may not recognize classes and resources added later. - `SplitCompatApplication` or `SplitCompat.install(this)` is required. - Without it, the app can encounter `ClassNotFoundException` or resource-loading failures. ### Native SO Libraries - DFM native libraries are installed in SplitCompat-managed paths rather than the standard `nativeLibraryDir`. - `SplitInstallHelper.loadLibrary()` may be required instead of `System.loadLibrary()`. - Unmodifiable third-party SDKs may need to remain in the base module. - Base and DFM modules must use identical ABI filters, or bundle creation can fail. ### STL and Native Version Conflicts - Multiple native libraries may share `libc++_shared.so`. - Because Android’s linker follows a “first-loaded wins” rule, the base module’s STL version can be used by DFM libraries. - NDK or STL mismatches can cause symbol conflicts, runtime crashes, or memory corruption. - The team used packaging exclusions and `pickFirsts` rules to control which libraries were included in each module. ### R8 Configuration - Keeping R8 rules only in the DFM caused base-module types to be obfuscated incorrectly. - The team centralized keep rules in the base module and minimized DFM consumer rules. ## Final Architecture: Separate Only the SO Files - Most of the remaining 40MB came from native video-processing, filtering, and rendering libraries. - The editor’s Kotlin/Java code occupied only a few hundred kilobytes. - The team kept feature code in the base module and delivered only the large SO files on demand. ### Benefits - If the DFM download or native library loading fails, the app can show an error instead of crashing. - Local development and debugging remain straightforward. - Separating only native libraries still reduced the module size by more than 90%. ### Trade-offs - Engineers must explicitly manage module installation, library loading, and feature startup. - Keeping code in the base while native binaries reside in the DFM is less intuitive. - The launch flow checks installation, loads the library, and then starts the editor. ## Testing and Operational Practices - Test both installed and uninstalled DFM states. - Validate real download behavior through Play Console internal testing or Internal App Sharing. - Use `bundletool` for local iteration and device-specific APK generation. - Universal APKs can simplify rapid installation and verification. - Test module installation and failure scenarios, not just normal editor execution. The main recommendation is to optimize for stability and maintainability rather than complete architectural separation. For large Android apps, isolating the largest native dependencies with on-demand DFM delivery can protect global users from unnecessary downloads while preserving team autonomy and operational simplicity.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Toxic combinations: when small signals add up to a security incident

Small security signals can become dangerous when combined: automated bots probing sensitive paths, unusual request behavior, and weak authentication or configuration. Cloudflare argues that analyzing these signals together—rather than judging each request independently—can reveal likely attack campaigns before compromise. Although toxic combinations are uncommon outside WordPress, the affected hosts may be highly exposed. ## What “Toxic Combinations” Mean - A toxic combination occurs when attackers compound several minor weaknesses into a viable breach. - Relevant signals include: - Bot activity and automated scanning - Sensitive paths such as `/admin`, `/debug`, `/metrics`, search, and payment endpoints - Anomalies such as unexpected HTTP status codes, geographic jumps, identity mismatches, high identifier churn, distributed rate-limit evasion, and traffic spikes - Missing session cookies or authorization headers and predictable identifiers - Traditional WAF, bot, and API defenses often assess the risk of individual requests. - Cloudflare’s approach examines the broader context across multiple requests, hosts, and paths. ## Measuring Exposure - Cloudflare analyzed a 24-hour sample of application-security data. - About 11% of analyzed hosts appeared susceptible to toxic combinations, largely because of vulnerable WordPress sites. - Excluding WordPress, only about 0.25% of hosts showed signs of exploitable combinations. - The analysis separated attacks into three stages: - **Hosts probed:** systems receiving requests for sensitive paths such as `/wp-admin` - **Hosts matching a toxic combination:** systems meeting the full detection criteria - **Reachable hosts:** systems that successfully responded to an exploit attempt - A `200 OK` response alone is not proof of exposure. Cloudflare recommends validating results against authentication requirements, redirects, and origin configurations to eliminate false positives. ## Probing Administrative Endpoints - Automated scanners targeted common administrative interfaces, including: - WordPress `/wp-admin` pages - Database management tools - Server dashboards - Cloudflare’s Log Explorer query groups successful requests by host, filters for likely bot traffic using a low bot score, and searches for configurable path patterns. - The query also excludes hosts represented only by raw IP addresses unless that filter is removed. ## Why Public Admin Panels Are Dangerous - Exposed administrative panels enable brute-force login attempts. - A successful compromise can allow attackers to: - Identify software and versions such as WordPress or Tomcat - Search for relevant CVEs and launch targeted exploits - Add the compromised host to a botnet that scans other websites - A sensitive endpoint returning successfully should therefore be tested for actual reachability and authentication weakness, not treated as conclusive evidence on its own. ## Practical Recommendation Monitor combinations of bot activity, sensitive-path access, anomalous behavior, and missing authentication signals. Investigate confirmed reachable endpoints, restrict or protect administrative interfaces, remove debug exposure, and validate detection queries against real application behavior to distinguish exploitable systems from false positives.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Bringing more transparency to post-quantum usage, encrypted messaging, and routing security

Cloudflare Radar is expanding its security coverage with new visibility into post-quantum encryption, Key Transparency for encrypted messaging, and ASPA deployment for routing security. The updates extend monitoring from user-to-Cloudflare connections to origin servers, provide tools for testing individual websites, and expose verification data that users can independently inspect. Together, they aim to make emerging Internet security technologies more measurable and transparent. ## Measuring Origin Post-Quantum Support - Cloudflare has tracked browser and client support for post-quantum encryption since 2024, rising from below 3% to more than 60% by February 2026. - The monitored algorithm, `X25519MLKEM768`, combines: - Classical X25519 key exchange - NIST-standardized ML-KEM post-quantum cryptography - Radar now measures whether customer origin servers support the same hybrid key exchange. - Cloudflare’s automated TLS scanner probes TLS 1.3-compatible origins and aggregates results daily. - The data measures algorithm support, not necessarily algorithm preference; a server’s TLS configuration can still choose a classical exchange even when post-quantum support exists. - Approximately 10% of origins currently support post-quantum-preferred key agreement, up from less than 1% in early 2025. - Adoption has accelerated as newer versions of OpenSSL, GnuTLS, and Go enabled hybrid post-quantum support by default. - Origin readiness data is available through Radar, Data Explorer, and the Radar API. ## Website Post-Quantum Compatibility Testing - Radar now includes a tool for testing whether a publicly accessible hostname supports post-quantum encryption. - Users can enter a hostname and optionally specify a port, with HTTPS port 443 used by default. - Results show: - Whether the connection is post-quantum secure - The negotiated TLS key exchange algorithm - The tool uses Cloudflare Containers to run a Go-based TLS scanner. - Because Workers cannot inspect the underlying TLS handshake, the container uses Go’s `crypto/tls` package to perform the connection and report the negotiated algorithm. - Cloudflare has consolidated its client- and origin-facing post-quantum measurements into a dedicated Radar section. ## Key Transparency for Encrypted Messaging - End-to-end encrypted services such as WhatsApp and Signal depend on correct public-key distribution. - If a messaging provider’s key database were compromised, an attacker could replace a contact’s public key and potentially intercept messages without detection. - Key Transparency mitigates this risk through an auditable, append-only public-key log. - The model is comparable to Certificate Transparency: - Messaging services publish users’ public keys to a transparency log. - Independent auditors verify that the log is correctly built and remains consistent. - Radar now provides a public dashboard for Key Transparency Logs used by E2EE messaging services. - The dashboard shows when each log was last signed and verified by Cloudflare’s Auditor. - Users can also access an API to independently validate the Auditor’s proofs. ## Routing Security and ASPA - Radar’s routing security coverage now includes global, country-level, and network-level information about ASPA deployment. - ASPA is an emerging standard intended to help detect and prevent BGP route leaks. - The new data extends Radar’s broader monitoring of Internet routing security. Cloudflare’s additions make post-quantum readiness, encrypted-message key integrity, and routing protection easier to measure and verify. Organizations can use the Radar dashboards, API, and hostname testing tool to assess their own migration and security posture.

Read original(opens in new tab)
cloudflare3 min readCurated summary

ASPA: making Internet routing more secure

ASPA (Autonomous System Provider Authorization) extends RPKI-based routing security from verifying a route’s destination to validating the path it takes. By publishing cryptographically signed lists of authorized providers, networks can detect BGP route leaks and some forged-origin hijacks. Cloudflare Radar now tracks ASPA adoption and records across the five Regional Internet Registries. ## From Origin Validation to Path Validation - RPKI uses Route Origin Authorizations (ROAs) to verify that an Autonomous System (AS) is authorized to announce specific IP prefixes. - ROAs protect against origin hijacks, where an attacker falsely claims ownership of someone else’s address space. - ASPA complements ROAs by validating the AS_PATH—the sequence of networks through which a route propagates. - Each AS publishes its authorized upstream providers, allowing other networks to check whether the observed path follows approved relationships. ## Detecting Route Leaks - Normal Internet routing generally follows a “valley-free” pattern: - Traffic moves upward from a customer through providers. - It may cross a peering connection near the top. - It then moves downward through providers toward the destination. - A route leak creates a “valley,” such as traffic traveling down to a customer and then back up to another provider. - ASPA evaluates the path from both directions: - The “up-ramp” is checked from the route origin toward its providers. - The “down-ramp” is checked backward from the destination. - A path is valid when both authorized chains meet. If they leave a gap, the route is considered ASPA Invalid, indicating a likely leak or unauthorized propagation. ## Example of ASPA Validation - In the example, AS65539 receives a route from customer AS65538. - AS65538 improperly propagates traffic received from provider AS65537 toward another provider, acting as a bridge between providers. - The upstream validation chain ends at AS65537, while the downstream chain ends at AS65538. - Because the two chains do not connect, ASPA identifies the route as invalid. ## Protection Against Forged-Origin Hijacks - ASPA can detect attacks in which the legitimate origin AS is retained but the attacker fabricates the path leading to it. - The victim’s signed provider list reveals that the attacker is not an authorized provider, allowing the route to be rejected. - ASPA is not universal protection: the article notes that a provider may still forge a path advertisement to its customer, a case the supplied text does not fully explain. ## Monitoring ASPA Adoption - Cloudflare Radar’s ASPA deployment monitoring feature shows adoption trends across all five RIRs. - It also lets users inspect ASPA records and changes for individual Autonomous Systems. ASPA provides an important second layer of routing security: ROAs verify where traffic should end, while ASPA helps verify how it gets there. Wider publication and validation of ASPA records should make route leaks and path manipulation easier to detect and prevent.

Read original(opens in new tab)
cloudflare3 min readCurated summary

We deserve a better streams API for JavaScript

Web Streams established a cross-runtime standard for handling streaming data, but their design reflects constraints from 2014–2016 rather than modern JavaScript practices. James M. Snell argues that the API’s reader, lock, and controller machinery creates unnecessary complexity and performance costs. He presents an alternative based on JavaScript language primitives that reportedly runs 2× to 120× faster across browsers and major runtimes. ## Historical Design Constraints - The WHATWG Streams Standard aimed to provide portable APIs for creating, composing, and consuming streams. - It was adopted by browsers, Cloudflare Workers, Node.js, Deno, Bun, and APIs such as `fetch()`. - The design predates JavaScript async iteration, which was standardized in ES2018. - Because `for await...of` did not yet exist, Web Streams introduced a separate reader/writer acquisition model. ## Excessive Ceremony for Basic Reads - Reading a stream to completion traditionally requires: - Calling `stream.getReader()`. - Repeatedly awaiting `reader.read()`. - Checking `{ value, done }` on every iteration. - Releasing the reader lock in a `finally` block. - These steps are API choices rather than inherent requirements of streaming. - Modern async iteration reduces the same operation to: ```js for await (const chunk of stream) { chunks.push(chunk); } ``` - However, async iteration was added after the original design, so it does not eliminate the underlying reader, lock, and controller complexity. - Advanced features such as BYOB reads still require developers to use the lower-level APIs. ## Problems with Manual Locking - Calling `getReader()` places an exclusive lock on the stream. - While locked, other code cannot read, pipe, or cancel the stream directly. - Forgetting `reader.releaseLock()` can permanently prevent later consumers from using the stream. - The `locked` property indicates that a lock exists, but not who owns it, why it exists, or whether the reader remains usable. - Internal operations such as piping also acquire locks, which can make stream behavior surprising. - Lock-release behavior with pending reads was historically unclear and varied between implementations before being clarified by the specification. - Async iterables improve the user experience by handling reader and lock management automatically, but the underlying model remains complex. ## Proposed Direction - The post argues that Web Streams’ limitations are fundamental design consequences, not isolated bugs easily fixed through incremental changes. - A better API should be built around modern JavaScript primitives, especially async iteration. - The author’s alternative reportedly achieves between 2× and 120× the performance of Web Streams across Cloudflare Workers, Node.js, Deno, Bun, and major browsers. - The claimed gains come from different architectural choices rather than narrowly optimized implementations. A more modern streams API should make common operations natural, avoid exposing fragile manual lock management, and use JavaScript’s native asynchronous iteration model from the start.

Read original(opens in new tab)
cloudflare3 min readCurated summary

The most-seen UI on the Internet? Redesigning Turnstile and Challenge Pages

Cloudflare redesigned Turnstile and Challenge Pages because these security interfaces are encountered billions of times daily and increasingly interrupt users as bot attacks grow. The redesign focused on reducing frustration through consistent information architecture, clearer language, better accessibility, and a deeper understanding of user journeys. The central conclusion is that security products must be designed not only to stop bots, but also to provide a humane, understandable experience for people at global scale. ## A Security Interface Seen Everywhere - Turnstile and Challenge Pages are served approximately **7.67 billion times per day**. - Their enormous reach creates a responsibility to support users across: - Different languages and cultures - A wide range of technical abilities - Different ages and accessibility needs - Varying devices, network conditions, and environments - As bot attacks increase, users are encountering verification challenges more frequently: - **2023:** 2.14 billion daily checks - **2024:** 3 billion - **2025:** 5.35 billion - This represented an average year-over-year increase of **58.1%**, making usability increasingly important. ## Auditing the Existing Experience Cloudflare reviewed every state, error message, and interaction in both products. - The audit found no consistent approach to error handling. - Some messages were overly technical and verbose, such as explanations involving incorrect device clocks or cached challenge pages. - Other messages were too vague, such as simply saying “Timed out.” - Layouts, visual hierarchy, and tone varied substantially between states. - User feedback mechanisms used ambiguous options like: - “The widget sometimes fails” - “The widget fails all the time” - These choices required frustrated users to interpret unclear distinctions and produced less useful feedback. - Challenge Pages also contained confusing states, technical jargon, and insufficient guidance about what users should do next. ## Mapping the Complete User Journey The team mapped both successful and unsuccessful paths through the verification experience. - The process covered initial encounters, errors, retries, and escalating frustration. - Designers collaborated with engineers who understood technical edge cases and product specialists who tracked user sentiment. - The team emphasized that technical sophistication does not automatically produce clear communication. - Interfaces needed to work for people with different: - Physical and mental capabilities - Cultural backgrounds - Ages - Levels of technical knowledge - At Cloudflare’s scale, unusual cases are common enough that they cannot be treated as negligible edge cases. ## Establishing a Unified Information Architecture Cloudflare applied the principle from *Don’t Make Me Think*: every moment users spend interpreting an interface creates friction, especially when they are already frustrated. - Previously, Turnstile and Challenge Pages placed information differently across states. - Users had to relearn where to find explanations, actions, and documentation links. - The redesign introduced one shared structure for both products. - Each experience would use: - The same visual hierarchy - Consistent placement for explanatory text - Consistent locations for actions - Consistent placement of documentation links - This approach limited some creative design options, but the team viewed those constraints as useful for improving clarity and consistency. Cloudflare’s redesign treats verification as a human-facing product rather than merely a security mechanism. A consistent structure, clearer messaging, and attention to accessibility can reduce the unnecessary frustration caused by challenges while preserving their protective purpose.

Read original(opens in new tab)
figma3 min readCurated summary

Workflow Lab: AI Image Tooling and Interactive Prototyping in Figma | Figma Blog

Figma’s Workflow Lab demonstrates how teams can combine AI image editing, vector tools, and interactive prototyping to test conversion-focused product experiences quickly. For Trivet, a recipe app, the goal is to increase referral-program signups without disrupting the user journey. The team explores three approaches—an immersive modal, an integrated feed card, and a full-screen overlay—each using different visual strategies to balance attention and usability. ## The Signup Problem - Trivet’s referral banner is generating fewer clicks and signups than expected. - The team must determine whether the banner is: - Too subtle to attract attention - Too disruptive to support a smooth experience - The sprint tests three levels of visibility: - A modal - An in-feed card - A full-screen overlay - The team evaluates not only layout, but also tone, visual tension, and how attention affects conversion. ## Direction One: Use Visual Treatment to Spark Curiosity - A modal creates a more immersive interruption than an in-feed card. - Photography is used to support messaging about “unlocking” exclusive recipes. - Recipes are partially blurred so they appear visible but inaccessible. - Figma’s glass effect, refraction, and progressive blur create a frosted, layered surface. - The visual treatment makes the reward feel tangible while justifying the interruption. ## Direction Two: Add Personality to the Interface - The referral message is integrated directly into the feed to feel more supportive and less disruptive. - A hand-drawn cake illustration replaces photography because it requires less visual space. - Figma’s Remove background and Vectorize tools turn the raster sketch into an editable vector asset. - The designer: - Adjusts colors using brand variables found through inline fill search - Uses the Cut tool to refine edges while preserving paths - Rotates anchor points with bounding boxes to polish details - The result is a reusable, scalable illustration that adds texture and approachability to the card. ## Direction Three: Create a Full-Screen Overlay - A full-screen experience maximizes visibility and clearly signals priority. - The larger format requires photography strong enough to carry the entire interaction. - Instead of replacing an imperfect image, the designer adapts it to the layout. - The workflow uses precise image-editing tools such as Erase object to remove distracting elements, including a spoon. - This approach prioritizes visual impact, though it introduces the greatest interruption to the user experience. ## From Design Exploration to Prototyping - The workflow connects Figma Design, AI image tools, FigJam, and Figma Make. - Designers can move from visual experimentation to interactive prototypes without waiting for production assets. - A broader team—including product, content, growth, and engineering—can compare the concepts and evaluate which level of attention best supports signups. The practical recommendation is to test visual intensity against user experience rather than assuming maximum visibility will perform best. Figma’s integrated image and prototyping tools make it faster to create and compare multiple directions before committing to implementation.

Read original(opens in new tab)