Techlist.io - Korean Tech Blog Curator

airbnb3 min readCurated summary

My Journey to Airbnb — Anna Sulkina

Anna Sulkina’s career journey moved from hardware diagnostics and frontend development into backend infrastructure and engineering leadership. Her experiences at Twitter taught her to design distributed systems for failure and to build consensus around transformative technologies like GraphQL. She joined Airbnb in 2022 because it aligned her passion for travel with an opportunity to strengthen developer infrastructure, organizational strategy, and engineering collaboration. ## Discovering Technology in Post-Soviet Ukraine - Sulkina grew up in Eastern Ukraine as the Soviet Union collapsed. - Her older brother introduced her to computers by bringing home hardware components and assembling a machine that loaded programs from a cassette player. - Seeing how individual components formed a working system inspired her to pursue technology. ## Learning English While Building Technical Skills - She studied programming at a Ukrainian university before immigrating to the United States. - Although she understood written English and knew how to program, communicating in English was initially more difficult than learning programming languages. - She took ESL classes while studying C++ and Java through Berkeley Extension. - Her first job was in hardware diagnostics at a five-person company. - A language barrier caused her to run out of time on a technical interview, but an interviewer familiar with her Berkeley class gave her another opportunity. - She eventually transitioned from C++ to Java, which became her primary language for many years. ## Moving Down the Stack and Into Leadership - Sulkina’s career progressed from hardware diagnostics to frontend, backend, and infrastructure engineering. - At the same time, she increasingly took on leadership responsibilities. - At Caymas Systems, her manager recognized her leadership potential and showed her the difference effective leadership makes. - At Comcast, she moved from individual contributor to engineering manager. - Coaching engineers, building software collaboratively, and developing high-performing teams convinced her that leadership was the right path. ## Lessons from Twitter’s Distributed Systems - During nearly nine years at Twitter, Sulkina advanced from first-line manager to director. - She worked through major operational events, including the “fail whale” period and the tweetstorm surrounding Ellen DeGeneres’s viral selfie. - Twitter’s transition from a monolith to microservices taught her that failure is inevitable in complex systems. - Resilient distributed systems must be designed to handle failures rather than assuming failures can be prevented. - Her cultural lesson involved turning promising ideas into adopted technologies. - She helped bootstrap Twitter’s GraphQL API, replacing legacy REST services. - The effort required leadership support, cross-team consensus, and stakeholder alignment, but ultimately improved product teams’ development velocity. ## Choosing Airbnb - Airbnb contacted Sulkina in 2022, when she felt ready to move beyond a well-established organization at Twitter. - The company appealed to her because it combined her professional interests with her personal passion for travel; she had been an Airbnb guest since 2013. - Airbnb’s Developer Platform organization had strong work happening in separate silos but needed clearer strategy, direction, and trust across engineering. - Sulkina began by clarifying the organization’s purpose and future direction. - Her early priorities included strengthening the organization, coaching leaders, and creating alignment within the team and with the teams it supported. - Over the following years, this work produced a high-performing organization with clearer strategy, stronger execution, and a focus on delivering business value. Sulkina’s story emphasizes that technical growth, organizational leadership, and personal motivation can reinforce one another. Her experience suggests that successful engineering leaders design for failure, invest in alignment, and use clear strategy to turn fragmented efforts into meaningful platform-wide impact.

Read original(opens in new tab)
google3 min readCurated summary

Scheduling in a changing world: Maximizing throughput with time-varying capacity

The post presents scheduling algorithms for non-preemptive jobs when cloud capacity changes over time because of failures, maintenance, power limits, or higher-priority workloads. The goal is to maximize completed job value while respecting release times, deadlines, processing durations, and fluctuating parallel capacity. The research establishes the first constant-factor guarantees for several offline and online variants, including a 1/11 competitive ratio for a demanding common-deadline model. ## Scheduling with Time-Varying Capacity - A capacity profile specifies how many jobs can run simultaneously at each point in time. - Each job has: - A release time - A hard deadline - A processing duration - A weight or profit - Jobs must run continuously once started in the non-preemptive setting. - If capacity drops during execution, an interrupted job loses its progress. - The objective is to select and schedule jobs maximizing total completed weight. - The study considers: - **Offline scheduling**, where future jobs and capacity changes are known. - **Online scheduling**, where jobs arrive dynamically and decisions cannot be reversed. ## Offline Scheduling Results - The optimal problem is NP-hard, so the work focuses on approximation guarantees. - For unit-profit jobs, an earliest-finish-time Greedy algorithm achieves a **1/2-approximation**. - It completes at least half as many jobs as an optimal schedule. - This matches the classic guarantee for single-capacity scheduling. - For jobs with different weights, a primal-dual algorithm achieves a **1/4-approximation**. ## Why Online Non-Preemptive Scheduling Is Difficult - Online schedulers must commit without knowing future jobs. - Starting a long job can block many shorter jobs that arrive later. - Because each completed job may have equal value regardless of duration, one poor decision can sharply reduce throughput. - Consequently, standard non-preemptive online algorithms have competitive ratios approaching zero. ## Interruption with Restarts - An active job may be interrupted, but its completed work is discarded and the job can be retried later. - A modified earliest-finish-time Greedy algorithm achieves a **1/2 competitive ratio**. - This means it can guarantee at least half the throughput of an optimal schedule with complete knowledge of future arrivals. ## Interruption Without Restarts - If an interrupted job is permanently discarded, online scheduling becomes substantially harder. - In general, every online algorithm can be forced into decisions that prevent it from completing much future work. - The competitive ratio again approaches zero. - The authors therefore study a practical special case in which all jobs share a common deadline. ## A Common-Deadline Algorithm For a unit-capacity system, the algorithm maintains a tentative schedule of jobs in disjoint time intervals. When a new job arrives, it applies the first suitable action: 1. Place the job in an empty interval. 2. Replace a scheduled future job if the new job is significantly shorter. 3. Interrupt the current job if the new job is shorter than its remaining processing time. 4. Discard the new job. - The approach balances immediate execution against preserving capacity for shorter future jobs. - A generalized version works with arbitrary capacity profiles. - The resulting algorithm achieves the first constant competitive guarantee for this setting: **1/11**. The results suggest that schedulers for volatile cloud environments need controlled interruption and carefully designed replacement policies. Allowing restarts offers strong guarantees, while stricter interruption rules require additional structure—such as a shared deadline—to achieve predictable performance.

Read original(opens in new tab)
google3 min readCurated summary

Beyond one-on-one: Authoring, simulating, and testing dynamic human-AI group conversations

DialogLab is an open-source research prototype for designing, simulating, and evaluating dynamic human–AI group conversations. It addresses the tension between rigid scripts and unpredictable generative dialogue by combining structured conversational phases with real-time improvisation. Its evaluation with 14 participants suggests that human-guided simulation offers the strongest balance of realism, engagement, and control. ## A Framework for Multi-Party Conversations - DialogLab separates a conversation’s social structure from its progression over time. - **Group dynamics** define: - Groups, such as a conference or social event - Parties, such as presenters and audiences - Elements, including human or AI participants and shared content - **Conversation-flow dynamics** define: - Snippets, or distinct phases such as opening, debate, and consensus - Participants and turn sequences within each snippet - Interaction styles, including collaborative or argumentative modes - Rules for interruptions and backchanneling - This separation makes complex conversation designs modular and easier to revise. ## The Author–Test–Verify Workflow ### Authoring with Visual Tools - Designers use a drag-and-drop canvas to arrange avatars and shared content. - Inspector panels configure personas, roles, interaction patterns, and snippet behavior. - Automatically generated prompts can be customized for specific narrative or conversational goals. ### Human-in-the-Loop Simulation - A live preview displays the evolving transcript. - In human-control mode, an audit panel suggests possible AI responses. - Designers can edit, accept, or reject suggestions, retaining control over the agents’ contributions. - The system supports both structured interactions and more improvisational conversations. ### Verification and Analytics - A verification dashboard provides post-hoc analysis of the conversation. - Visualizations show turn-taking distributions and sentiment flows. - These tools help creators diagnose interaction patterns without manually reviewing entire transcripts. ## Prototype Evaluation - Fourteen participants from game design, education, and social science research evaluated DialogLab. - They designed an academic social event and tested AI group discussions under three conditions: - **Human control:** Users prompted agents to shift topics, introduce perspectives, ask probing questions, or generate emotional responses. - **Autonomous:** Agents participated proactively according to predefined random or sequential orders. - **Reactive:** A simulated human agent responded only when directly addressed. - Human control was rated significantly more engaging and was generally considered more effective and realistic. - Participants also described the interface as intuitive, flexible, and enjoyable. - Users valued the combination of automated prompt generation, detailed customization, and support for different moderation strategies. DialogLab demonstrates that effective multi-party conversational design benefits from combining explicit structure with controlled improvisation. For developers and researchers building group-based human–AI experiences, a visual authoring workflow paired with human-guided simulation and analytics can provide a practical foundation for rapid iteration and more realistic interactions.

Read original(opens in new tab)
gitlabOriginal article

DevSecOps-as-a-Service on Oracle Cloud Infrastructure by Data Intensity (opens in new tab)

Data Intensity’s DevSecOps-as-a-Service provides a solution for organizations that require the granular control of GitLab Self-Managed but wish to eliminate the operational burden of infrastructure maintenance. By hosting dedicated GitLab instances on Oracle Cloud Infrastructure (OCI), the service combines the security and customization of a self-managed environment with the convenience of a fully managed platform. This partnership enables teams to focus on software delivery while leveraging expert management for high availability and disaster recovery. ### The Benefits of GitLab Self-Managed * Offers complete ownership of data residency and instance configuration to meet strict regulatory and compliance requirements. * Enables deep customization and integration possibilities that are often restricted in standard SaaS environments. * Addresses the challenges of manual server management, upgrades, and high-availability scaling by offloading these tasks to a managed provider. ### Managed Service Features and Support * Provides 24/7 monitoring, alarming, and expert technical support for standalone GitLab instances. * Includes scheduled quarterly patching performed during customer-specified maintenance windows to minimize disruption. * Ensures business continuity through automated backups and professional disaster recovery protection. * Utilizes tiered architectures designed to scale based on specific user capacities and recovery time objectives. ### Infrastructure Optimization via OCI * Delivers significant cost efficiency, with organizations typically realizing 40-50% reductions in infrastructure spending compared to other hyperscalers. * Supports diverse deployment models, including Public Cloud, Government Cloud, EU Sovereign Clouds, and dedicated infrastructure behind a corporate firewall. * Maintains consistent pricing and operational tooling across hybrid, global, and regulated environments. ### Implementation and Migration * Data Intensity offers optional migration services to transition existing code repositories and configurations to the OCI environment seamlessly. * The service is specifically designed for organizations with predictable cost requirements and those lacking in-house infrastructure expertise. * Deployment planning involves tailored consultations to match specific compliance and data residency needs with OCI’s global region availability. This managed service is a recommended path for enterprise teams that need to prioritize data sovereignty and flexibility without sacrificing the speed of a turnkey solution. Organizations currently using or planning to adopt OCI can leverage this service to standardize their DevSecOps workflows while achieving significant infrastructure savings.

grammarlyOriginal article

Grammarly’s AI Detector Agent Ranks #1 in Quality (opens in new tab)

Grammarly has launched a high-ranking AI detection tool specifically designed for students and educational institutions to address the growing complexity of machine-generated content. By integrating this detector into their existing ecosystem, the company aims to provide a reliable way to verify human authorship while protecting the integrity of a student's original voice. ### Implementing Reliable AI Detection (RAID) * Grammarly utilizes the RAID (Reliable AI Detection) framework to ensure the tool remains effective against evolving large language models (LLMs). * The detector focuses on minimizing false positives, which is critical in academic settings to avoid wrongful accusations of misconduct. * The system is benchmarked to provide high-performance accuracy, offering institutions a standardized metric for evaluating the authenticity of submitted work. ### Preserving Human Authorship and Voice * The widespread use of generative AI has created a climate of skepticism where students’ original work is frequently questioned by instructors and automated systems. * The detector provides a nuanced analysis that helps distinguish between legitimate AI-assisted refinement—such as grammar and clarity checks—and full AI content generation. * By offering transparent reporting, the tool helps students validate their personal writing process and defend the originality of their voice. ### Multi-Agent Integration and Ecosystem Support * AI detection is positioned as a single "agent" within a broader suite of writing, editing, and citation tools. * The tool is built to integrate seamlessly with institutional workflows and Learning Management Systems (LMS), ensuring it is accessible at the point of writing. * This holistic approach treats detection as part of a supportive writing environment rather than a punitive standalone feature, encouraging responsible AI use. To maintain trust in digital communication, institutions should adopt detection tools that prioritize reliability and transparency, ensuring that the transition to AI-integrated learning does not come at the expense of student confidence or academic honesty.

aws3 min readCurated summary

AWS Weekly Roundup: Claude Opus 4.6 in Amazon Bedrock, AWS Builder ID Sign in with Apple, and more (February 9, 2026) | Amazon Web Services

The February 9, 2026 AWS roundup highlights updates across infrastructure, security, databases, and AI. Major announcements include new EC2 instances, cross-account DynamoDB replication, improved identity controls, CloudFront mutual TLS, Claude Opus 4.6 in Bedrock, and structured model outputs. AWS also announced AWS Community Day Romania for April 23–24, 2026. ## Compute, Networking, and Configuration - **New EC2 C8id, M8id, and R8id instances** - Powered by custom Intel Xeon 6 processors. - Deliver up to 43% higher performance and 3.3× more memory bandwidth than previous-generation instances. - **AWS Network Firewall price reductions** - Reduces hourly and data-processing costs for NAT Gateways service-chained with Network Firewall secondary endpoints. - Removes additional data-processing charges for Advanced Inspection and TLS inspection. - **Amazon ECS Network Load Balancer support** - Enables managed linear and canary deployments for applications using NLBs. - Supports TCP/UDP workloads, low-latency services, long-lived connections, and static IP requirements. - **Expanded AWS Config coverage** - Adds support for 30 resource types across services such as Amazon EKS, Amazon Q, and AWS IoT. - Improves resource discovery, auditing, assessment, and remediation. ## Databases and Operations - **Cross-account DynamoDB global table replication** - Allows multi-Region, multi-active tables to replicate across AWS accounts. - Improves resilience, account-level workload isolation, and independent security and governance controls. - **Improved Amazon RDS connection experience** - Generates connection snippets for Java, Python, Node.js, `psql`, and other tools. - Adjusts examples automatically for authentication settings, including IAM token-based authentication. - Adds CloudShell integration for connecting to databases directly from the RDS console. ## Identity and Security - **AWS Builder ID adds Sign in with Apple** - Apple users can access services such as AWS Builder Center, Training and Certification, re:Post, AWS Startups, and Kiro. - Complements the existing Google sign-in option. - **More identity-provider claim validation in AWS STS** - Supports selected claims from Google, GitHub, CircleCI, and OCI. - These claims can be used as condition keys in IAM trust policies and resource control policies for more precise federated-access controls and data perimeters. - **Account names in the AWS Management Console** - Displays the account name in the navigation bar, making it easier to distinguish between authorized AWS accounts. - **CloudFront origin mutual TLS** - Lets CloudFront authenticate to origins using certificates. - Helps restrict backend access to verified CloudFront distributions across AWS, on-premises, third-party cloud, and external CDN environments. ## AI and Amazon Bedrock - **Claude Opus 4.6 available in Amazon Bedrock** - Anthropic’s latest model targets complex coding, agentic tasks, enterprise workflows, and professional work requiring deep reasoning and reliability. - **Structured outputs in Amazon Bedrock** - Models can return responses matching developer-defined JSON schemas. - Reduces the need for prompt-based JSON enforcement and additional validation, making production integrations more predictable. ## Upcoming AWS Event - **AWS Community Day Romania — April 23–24, 2026** - Features more than 10 technical sessions from AWS Heroes, Solutions Architects, and industry experts. - Includes networking opportunities for developers, architects, entrepreneurs, and students. These updates emphasize stronger infrastructure performance, better multi-account governance, more secure authentication, and more reliable AI application development. Teams should evaluate the new services based on their networking, resiliency, identity, and structured-output requirements.

Read original(opens in new tab)
meta3 min readCurated summary

Building Prometheus: How Backend Aggregation Enables Gigawatt-Scale AI Clusters

Meta uses Backend Aggregation (BAG) as a high-capacity Ethernet super-spine to connect tens of thousands of GPUs across data centers and regions. In the Prometheus AI cluster, BAG links regional networks and Meta’s backbone while bridging two L2 fabric technologies: Disaggregated Schedule Fabric (DSF) and Non-Scheduled Fabric (NSF). Its modular hardware, resilient topologies, and advanced routing are designed to deliver reliable, petabit-scale connectivity for a gigawatt-scale AI system. ## Backend Aggregation’s Role - BAG interconnects multiple spine fabrics across data centers and regions. - It aggregates regional networks and connects them to Meta’s backbone. - Inter-BAG capacity can reach 16–48 petabits per second per regional pair. - Prometheus will span multiple buildings and connect tens of thousands of GPUs. ## Regional BAG Connectivity - BAG layers are distributed regionally to serve groups of L2 fabrics while respecting distance, latency, and buffer constraints. - Two connection topologies are used: - **Planar topology:** One-to-one connections between BAG switches in different regions; simpler to manage but creates more concentrated failure domains. - **Spread topology:** Links are distributed across switches and planes, improving path diversity and resilience. - The choice depends on site size and available fiber. ## Connecting DSF and NSF Fabrics - Meta’s L2 networks use both: - **Disaggregated Schedule Fabric (DSF)** - **Non-Scheduled Fabric (NSF)** - DSF zones across multiple buildings connect to BAG through backend edge pods. - NSF connects to BAG planes through matching Spine Training Switches. - Oversubscription is carefully managed: - L2-to-BAG oversubscription is typically around 4.5:1. - One NSF example has an effective ratio of 4.98:1. - BAG-to-BAG ratios vary by region and link capacity. ## Hardware and Routing - BAG uses modular chassis with Jericho3 ASIC line cards. - Each line card supports up to 432 800G ports. - Larger central-hub chassis support many spoke connections and long-distance links. - eBGP with link-bandwidth attributes enables Unequal Cost Multipath (UCMP), improving load balancing and failure recovery. - BAG-to-BAG links use MACsec for network security. ## Resilience and Failure Management - The design includes detailed port striping, IP addressing, and failure-domain analysis. - Failures are evaluated at the BAG, data-hall, and power-distribution levels. - Mitigation techniques include: - Draining affected BAG planes - Conditional route aggregation - Reducing blackholing risks during failures ## Managing Long-Distance Links - Distributed BAG architecture keeps L2-to-edge distances short, which benefits shallow-buffer NSF switches. - Longer BAG-to-BAG connections require deep-buffer switches. - These buffers provide headroom for lossless congestion-control mechanisms such as Priority Flow Control (PFC). ## Broader Impact BAG provides the networking foundation for Prometheus and future AI clusters. By combining regional aggregation, high-density hardware, resilient connection topologies, and fabric interoperability, Meta can scale AI infrastructure across multiple data centers while maintaining bandwidth, reliability, and operational flexibility.

Read original(opens in new tab)
tossOriginal article

From Perimeter Security to Zero (opens in new tab)

Toss Payments transformed its security infrastructure from a vulnerable, single-layered legacy system into a robust "Defense in Depth" architecture spanning hybrid IDC and AWS environments. By integrating advanced perimeter defense, internal server monitoring, and container runtime security, the team established a comprehensive framework that prioritizes visibility and continuous verification. This four-year journey demonstrates that modern security requires moving beyond simple boundary protection toward a proactive, multi-layered strategy that assumes breaches can occur. ### Perimeter Defense and SSL/TLS Visibility * Addressed the critical visibility gap in legacy systems by implementing dedicated SSL/TLS decryption tools, allowing the team to analyze encrypted traffic for hidden malicious payloads. * Established a hybrid security architecture using a combination of physical DDoS protection, IPS, and WAF in IDC environments, complemented by AWS WAF and AI-based GuardDuty in the cloud. * Developed a collaborative merchant response process that moves beyond simple IP blocking; the system automatically detects malicious traffic from partners and provides them with detailed vulnerability reports and remediation guides (e.g., specific SQL injection points). ### Internal Network Security and "Assume Breach" Monitoring * Implemented **Wazuh**, an open-source security platform, in IDC environments to monitor lateral movement, collect centralized logs, and perform file integrity checks across diverse operating systems. * Leveraged **AWS GuardDuty** for intelligent threat detection in the cloud, focusing on malware scanning for EC2 instances and monitoring for suspicious process activities. * Established automated detection for privilege escalation and unauthorized access to sensitive system files, such as tracking instances where root privileges are obtained to modify the `/etc/passwd` file. ### Container Runtime Security as the Final Defense * Adopted **Falco**, a CNCF-hosted runtime security tool, to protect Kubernetes environments by monitoring system calls (syscalls) in real-time. * Configured specific security rules to detect "container escape" attempts, unauthorized access to sensitive files like `/etc/shadow`, and the execution of new or suspicious binaries within running containers. * Integrated **Falco Sidekick** to manage security events efficiently, ensuring that anomalous behaviors at the container level are instantly routed to the security team for response. ### Zero Trust and Continuous Verification * Shifted toward a Zero Trust model for the internal work network to ensure that all users and devices are continuously verified regardless of their location. * Focused on implementing dynamic access control and the principle of least privilege to minimize the potential impact of credential theft or device compromise. Organizations operating in hybrid cloud environments should move away from relying on a single perimeter and instead adopt a multi-layered defense strategy. True security resilience is achieved by gaining deep visibility into encrypted traffic and maintaining granular monitoring at the server and container levels to intercept threats that inevitably bypass initial defenses.

figma3 min readCurated summary

Why Demand for Designers Is on the Rise | Figma Blog

Companies are investing more in design, with 82% of leaders reporting that demand has either increased or remained steady. The post argues that AI is not reducing the need for designers; instead, it is increasing demand for people who can use AI tools, design AI products, and connect design with strategy and business growth. Fast-growing organizations are leading this hiring momentum, while employers increasingly favor experienced, AI-fluent candidates. ## Design hiring is increasing across industries - Nearly half of hiring managers say demand for designers has increased, and most of them report growth of at least 10%; more than a quarter report increases of 25% or more. - Technology companies lead hiring, but demand is also growing in sectors such as retail, publishing, aviation, and other non-tech industries. - Companies are hiring designers to improve digital experiences, strengthen online presence, and create new customer value. - Planned hiring varies by company growth: - 46% of fast-growing companies expect to increase hiring. - 40% of average-growth companies plan to do so. - 33% of slower-growth companies expect increased hiring. - High-growth companies view design as a way to test ideas earlier, move faster, differentiate products, and drive revenue. - Although only 20% of managers believe the overall hiring market is improving, 40% plan to add design headcount within six months. - Design job postings among Designer Fund portfolio companies reportedly rose about 60% in 2025 compared with 2024. ## AI is fueling demand for designers - Rapid advances in AI models and tools are creating demand for designers who can immediately work with evolving AI processes. - Employers want both: - Proficiency with AI tools in everyday design workflows. - Experience designing AI-powered products. - 73% of hiring managers report an increasing need for AI-tool proficiency. - 79% report an increasing need for knowledge of designing AI products. - AI fluency is increasingly treated as a hiring requirement rather than an optional advantage. - Companies are prioritizing candidates who combine technical ability, strategic thinking, experimentation, and approaches such as human-in-the-loop and human-augmented AI. ## Seniority and broader judgment matter - The article begins a discussion of companies prioritizing senior talent, particularly as teams face pressure to deliver quickly. - Hiring managers are looking for designers with strong skills, judgment, and experience—not only executional ability. - The combination of design expertise, strategic thinking, and AI capability is becoming increasingly valuable. Designers can improve their prospects by developing practical AI fluency alongside core design skills, learning how to design AI products, and demonstrating strategic judgment and business impact.

Read original(opens in new tab)
google3 min readCurated summary

How AI trained on birds is surfacing underwater mysteries

Perch 2.0, Google DeepMind’s bioacoustics foundation model, was trained mainly on birds and terrestrial animals yet performs strongly on underwater audio. The study shows that its learned audio embeddings can support accurate whale, dolphin, reef-sound, and killer-whale classification with only a few labeled examples. This suggests that large, broadly trained bioacoustics models can transfer across environments and accelerate marine research without requiring extensive underwater training data. ## Underwater Mysteries and Bioacoustics - Ocean recordings reveal animal behavior, species distributions, and unexplained sounds. - The “biotwang,” recently attributed by NOAA to Bryde’s whales, illustrates how new calls and species identifications continue to emerge. - Google has previously developed models for humpback whales and multi-species whale detection. - Perch 2.0 extends this work despite having no underwater audio in its training data. ## Transfer Learning for Custom Classifiers - Researchers can use an existing model directly when its labels match their data. - For new sounds or datasets, transfer learning avoids training a deep neural network from scratch. - Perch 2.0 converts audio windows into compact numerical embeddings. - A logistic regression classifier is then trained on those embeddings using labeled examples. - This requires far less computation, experimentation, and training data than full neural-network training. ## Evaluation on Marine Datasets - The researchers tested Perch 2.0 with few-shot linear probes using 4, 8, 16, or 32 examples per class. - Performance was measured using ROC-AUC, where values closer to 1 indicate better class separation. - Evaluation datasets included: - **NOAA PIPAN:** Baleen-whale recordings, including minke, humpback, sei, blue, fin, and Bryde’s whales. - **ReefSet:** Reef biological sounds, fish, dolphins, anthropogenic noise, and waves. - **DCLDE:** Killer whales, humpbacks, abiotic sounds, unknown sounds, and killer-whale ecotypes. - More examples generally improved results. - ReefSet performance was already high with four examples per class for most models. - Perch 2.0 was consistently among the best-performing models across datasets and sample sizes. ## Comparisons with Other Models - Perch 2.0 was compared with Perch 1.0, SurfPerch, and Google’s multi-species whale model. - It also outperformed AVES-bird and AVES-bio on most underwater tasks. - The results show that strong underwater transfer is not limited to models trained on marine audio. ## Why Bird-Based Training Transfers to Whales - The authors suggest that large models trained on extensive datasets can generalize effectively to unfamiliar downstream tasks. - Shared acoustic patterns across animal vocalizations may allow representations learned from birds and other terrestrial species to remain useful underwater. - The findings challenge the assumption that a model must be trained directly on underwater recordings to perform well on marine classification tasks. ## Practical Tools for Researchers - Google provides a paper and a Google Colab tutorial. - The tutorial demonstrates an end-to-end workflow for building a whale-vocalization classifier. - It uses NOAA’s NCEI Passive Acoustic Data Archive and Google Cloud. - Researchers can create agile, task-specific models with relatively small labeled datasets. Perch 2.0 demonstrates that broad bioacoustic pretraining can substantially reduce the effort required to study marine sounds. Researchers can begin with general-purpose embeddings and adapt them to new whale species, calls, or underwater sound categories using only modest labeled data.

Read original(opens in new tab)
spotify3 min readCurated summary

Our Multi-Agent Architecture for Smarter Advertising | Spotify Engineering

The post argues that fragmented advertising workflows, not backend infrastructure, are the core problem. Although buying channels share services and data, their planning and optimization logic is repeatedly reimplemented across channels and surfaces, causing drift and technical debt. The proposed solution is a shared agentic decision layer that interprets advertiser goals, orchestrates existing Ads APIs, and applies consistent reasoning across products. ## Fragmented Workflows Across a Shared Backend - Direct, Self-Serve, and Programmatic buying use largely consolidated infrastructure but retain different workflows and decision logic. - Spotify Ads Manager, Salesforce, Slack, and internal tools contain overlapping automation. - Budget allocation, inventory selection, reach, efficiency, and STR decisions are repeatedly implemented in different places. - Incremental workflow changes therefore create duplicated maintenance work and inconsistent behavior. ## Why Conventional Workflow Services Fall Short - Hard-coded state machines and REST services are poorly suited to combinatorial planning tasks. - Campaign planning depends on: - User and advertiser characteristics - Available inventory and audiences - Business priorities - Forecasts, performance, and optimization goals - A workflow optimized for one channel or “happy path” will not adapt well as requirements change. - Improvements to decision logic must be replicated across every product surface, increasing the risk of divergence. ## The Missing Intent Layer - Existing systems can perform individual actions such as creating line items, running forecasts, and retrieving insights. - They do not consistently translate high-level objectives into: - A sequence of tool calls - Explicit tradeoffs - Validation and safety checks - An objective such as maximizing reach in Brazil while protecting video inventory and meeting STR requires coordinated reasoning across multiple capabilities. ## A Modular Agentic Architecture - Campaign planning and management are modeled as cooperating specialized agents. - Agents use shared signals, including: - Inventory - Audiences - STR - Quality and risk - Historical performance - They jointly optimize advertiser goals and Spotify’s business constraints. - Existing Ads services become tools that agents orchestrate, rather than capabilities being rebuilt in each workflow. - A long-running orchestration layer delegates tasks while agents share context and evaluation logic. - The same decision engine can support every buying channel and surface. ## Engineering Implications - APIs need to be designed as agent tools, rather than only as CRUD interfaces. - Testing must include behavioral evaluation in addition to unit and integration tests. - Observability should explain what an agent decided and why, not merely track latency and errors. - Safety requires guardrails for semi-autonomous decisions, beyond ordinary input validation. - The approach avoids both duplicated deterministic workflows and a brittle, centralized rules engine for probabilistic, ML-heavy advertising logic. The overall recommendation is to centralize campaign decision-making in a reusable agentic platform while keeping existing services as specialized tools. This should reduce duplicated workflow logic, make improvements consistent across products, and allow advertising workflows to evolve without repeatedly rebuilding them.

Read original(opens in new tab)
kakao4 min readCurated summary

In Search of Lost Reports: Kakao

KIMS, Kakao’s internal SMS platform, experienced rare cases where vendors sent delivery reports successfully, yet messages remained stuck in `SENT` instead of becoming `REPORTED`. The cause was a race condition: a fast vendor’s report arrived before the API server had committed the message record. The investigation showed that an unnecessarily long transaction—especially for paid messages with billing-event processing—delayed persistence and allowed valid reports to be dropped. ## KIMS Message Processing Flow - KIMS processes roughly one million SMS messages per day across multiple IDC environments and external vendors. - The normal flow is: - Route the request to a suitable vendor. - Call the vendor and record the message as `SENT`. - Deliver the message to the recipient. - Receive the vendor’s delivery report. - Update the message to `REPORTED`. - These stages run asynchronously across separate services, so their execution order is not guaranteed. ## Discovering the Missing Reports - Some messages remained in `SENT` even though Report Server logs confirmed that delivery reports had arrived. - The issue affected only about `0.02%` of messages, making it difficult to reproduce in tests or local environments. - Two patterns emerged: - Missing reports were concentrated among messages sent through one particular vendor. - Paid messages were affected more often than free messages. ## The Race Condition - The problematic vendor returned reports unusually quickly: - Other vendors typically took more than one second. - This vendor averaged around 20 ms. - Missing-report cases averaged only about 8 ms. - The API server performed additional processing before committing the message record. - For paid messages, billing-event publication was included in the same `@Transactional` scope, making the transaction longer. - Consequently, the sequence could become: 1. API Server calls the vendor. 2. API Server performs billing-related processing. 3. The vendor delivers the message and immediately sends a report. 4. Report Server receives the report before the message row exists in the database. 5. Report Server treats the report as invalid and drops it. 6. API Server finally commits the message as `SENT`. - The report was not lost at the network or vendor level; it was discarded because the system’s write path had not completed. ## Reducing Transaction Scope - The first fix was to remove nonessential work from the main transaction. - Billing-event publication was moved to asynchronous processing using `@Async` and `@TransactionalEventListener`. - The transaction was reduced to the essential state change and database commit. - This advanced the average commit point by approximately 10 ms and significantly reduced report omissions. - It also avoided a dual-write anti-pattern in which an external Kafka event was published inside a database transaction that could later roll back. ## Reconsidering the Need for a Transaction The incident prompted a broader review of whether the transaction was needed at all. - **Atomicity:** The transaction contained only one database write, with no multi-table or cross-record operation requiring all-or-nothing rollback. - **Read isolation:** Metadata such as vendor quality metrics was updated only every few minutes, and using a slightly stale value was acceptable. The independently read tables did not require a single consistent snapshot. - **Write isolation:** JPA’s dirty checking kept the status change in the persistence context until transaction completion, delaying the actual database write. This delay was precisely what allowed the report to arrive first. The article therefore presents the transaction itself—not the vendor or report receiver—as a source of unnecessary latency and an architectural anti-pattern in this workflow. ## Practical Recommendation Use transactions only when their guarantees are required. Keep critical persistence paths short, move external events and nonessential processing after commit, and critically evaluate whether delayed commit semantics could allow asynchronous consumers to observe a missing record.

Read original(opens in new tab)
line4 min readCurated summary

Creating the Cloud of the Future

LY Corporation is consolidating Yahoo! JAPAN and LINE’s internal cloud services into Flava, a private cloud for application development. The article outlines how Flava could evolve over the next two to three years through unified developer platforms, stronger yet more usable security, scalable multimedia storage, AI infrastructure, and intelligent cloud management. Its ultimate goal is to make complex infrastructure easier to consume while automating operational work. ## Platform Flavaization - Flava currently focuses on infrastructure, databases, and containers, while other development services are spread across separate internal platforms. - Developers must learn different systems for: - Access control and approvals - Logging, monitoring, metering, and billing - APIs, CLIs, and user interfaces - Multi-region and availability-zone operations - “Flavaization” means offering all development platforms through a consistent cloud experience. - LY expects much of this integration to be completed within the next one to two years. ## Stronger, More Usable Security - Flava incorporates security governance from the architecture and product-planning stages, working with the CISO organization. - Data environments are separated by security level: - Default - Secret - Top secret - Sensitive changes require role-based permissions, organizational reporting, expert review, and formal approval. - The main challenge is usability: - Resources can now be provisioned within minutes, but access may still require around ten workflows, such as VDI and Box account creation, taking up to two months. - VPC ACL controls can add several milliseconds of latency, which may affect latency-sensitive services such as LINE messaging. - Flava must provide “usable security” that preserves strong governance without making development excessively slow or difficult. ## Storage for Growing Multimedia Data - Users continuously generate and retain large volumes of photos, videos, and other multimedia content. - Storage demand can grow even when service traffic remains stable. - Flava needs storage technologies suited to different data lifecycles, balancing: - Cost - Throughput and latency - Searchability - Compression and deduplication - Encryption - Efficient tiered storage will be essential for managing long-lived user data economically. ## AI Operations Platforms - LY is adopting AI tools and agents across its organizations, creating demand for shared AIOps infrastructure. - Potential platform capabilities include: - Approved MCP server development and management - Vector databases - AI observability tools such as Langfuse - AI model management - Because AI systems handle internal data, these platforms must comply with company security and data-processing policies. - Flava aims to rapidly evaluate emerging AI technologies and provide compliant, standardized services across the company. ## Network and Storage Infrastructure for AI - AI workloads process larger datasets while requiring very low network latency and high throughput. - Relevant technologies include: - DPUs - Smart NICs - High-speed NVMe storage - Automated storage tiering - Operating networks and storage at cloud scale introduces major challenges in latency, reliability, fault tolerance, throughput, change management, and security. - Flava’s existing network and storage engineering teams have experience supporting LINE and Yahoo! JAPAN at large scale and will adapt that expertise for AI workloads. ## The Intelligent Cloud - Future users may describe infrastructure requirements in natural language rather than manually configuring resources through consoles, APIs, CLIs, or Terraform. - For example, Flava could translate requirements for image processing, AI-based content labeling, messaging, and tiered storage into an architecture and deployable system. - An intelligent Flava could also: - Generate network diagrams and ACL matrices - Identify vulnerabilities and prioritize remediation - Recommend cost optimizations - Detect underutilized resources - Find unencrypted personal information - Manage OSS vulnerability responses - Chatbots could automate tasks such as identifying low-utilization resources while excluding standby failover servers or proposing cost reductions for them. - Operational campaigns currently requiring substantial engineer participation could increasingly be handled by AI agents. Flava’s recommended direction is to combine a unified cloud experience with practical security, lifecycle-aware storage, AI-ready infrastructure, and natural-language automation. The article argues that building this future cloud requires both deep infrastructure expertise and strong attention to developer and user experience.

Read original(opens in new tab)
github3 min readCurated summary

Continuous AI in practice: What developers can automate today with agentic CI

Continuous AI extends CI into software-engineering tasks that require judgment, context, and interpretation rather than deterministic rules. It uses continuously running agents guided by natural-language instructions to review repositories, identify issues, and produce reviewable artifacts such as patches, issues, or reports. GitHub’s central argument is that AI should complement—not replace—traditional CI, while operating within explicit permissions and developer oversight. ## Why CI Isn’t Enough - CI is effective for binary, rule-based checks: - Tests pass or fail. - Builds succeed or fail. - Linters detect defined violations. - Many important engineering tasks depend on intent and context, including: - Finding discrepancies between documentation and implementation. - Detecting confusing accessibility text that passes linting. - Identifying behavioral changes caused by dependency updates. - Spotting subtle performance regressions, such as compiling a regular expression inside a loop. - Recognizing UI regressions that only appear during interaction. - GitHub describes this as a shift from AI-generated code toward AI handling cognitively demanding maintenance work. ## What Continuous AI Means - Continuous AI is a pattern, not a replacement for CI: - **Natural-language rules + agentic reasoning, executed continuously inside a repository.** - Developers describe expectations in natural language, especially when those expectations are difficult to encode with schemas, heuristics, or YAML. - Example workflows include: - Comparing documented behavior with implementation and proposing fixes. - Producing weekly reports on project activity, bug trends, and code churn. - Detecting performance regressions in critical paths. - Finding semantic regressions in user flows. - Workflows are refined collaboratively with agents by adding intent, constraints, and acceptable outputs rather than being authored as a perfect single instruction. ## Guardrails and Safe Outputs - Agents operate with read-only repository access by default. - They cannot modify content, create issues, or open pull requests unless explicitly authorized. - “Safe Outputs” defines the exact artifacts an agent may produce and the constraints governing them. - Agent activity is sanitized, logged, and auditable. - The goal is to keep the potential impact predictable even when agents make mistakes or behave unexpectedly. ## Natural Language Complements YAML - Deterministic problems should remain in CI, using YAML, schemas, tests, and heuristics. - Some expectations—such as determining whether documentation and code still express the same behavior—require semantic understanding. - Natural-language instructions let agents reason about intent without forcing that intent into brittle rules. - Continuous AI therefore expands automation into judgment-heavy tasks while preserving CI as the foundation for deterministic validation. ## Developers Remain in the Loop - Agents do not make unrestricted autonomous commits. - Depending on permissions, they can produce pull requests, issues, comments, discussions, or other reviewable artifacts. - Pull requests are especially useful because they fit existing developer review and collaboration practices. - The broader vision is to delegate recurring maintenance work while allowing developers to retain judgment, taste, and final control. Continuous AI is best adopted alongside traditional CI: use conventional automation wherever rules are sufficient, and use guarded, continuously running agents for tasks involving interpretation, synthesis, and evolving intent.

Read original(opens in new tab)
cloudflare3 min readCurated summary

2025 Q4 DDoS threat report: A record-setting 31.4 Tbps attack caps a year of massive DDoS assaults

Cloudflare’s 2025 DDoS report describes a dramatic escalation in both attack frequency and scale. DDoS attacks more than doubled to 47.1 million, while botnets such as Aisuru-Kimwolf launched unprecedented HTTP floods, including a record 31.4 Tbps attack. Cloudflare concludes that autonomous, adaptive mitigation is increasingly essential as attacks grow more frequent, larger, and more sophisticated. ## Record Growth in DDoS Attacks - Cloudflare mitigated 47.1 million DDoS attacks in 2025, a 121% increase from 2024 and a 236% increase since 2023. - The network automatically mitigated an average of 5,376 attacks per hour: - 3,925 network-layer attacks - 1,451 HTTP attacks - In Q4 2025, attacks increased 31% from the previous quarter and 58% year over year. - Network-layer attacks accounted for 78% of Q4 activity. ## Network-Layer Attacks More Than Triple - Network-layer attacks rose from 11.4 million in 2024 to 34.4 million in 2025. - An 18-day campaign in Q1 generated approximately 13.5 million attacks against Cloudflare infrastructure and Magic Transit customers. - The campaign used multiple vectors, including: - SYN floods - Mirai-generated attacks - SSDP amplification - Cloudflare’s systems detected and mitigated the campaign automatically. ## The Aisuru-Kimwolf “Night Before Christmas” Campaign - Beginning December 19, 2025, the Aisuru-Kimwolf botnet attacked Cloudflare and its customers with HTTP floods exceeding 20 million requests per second. - The botnet is estimated to contain 1–4 million malware-infected devices, primarily Android TVs. - During the campaign, Cloudflare mitigated 902 hyper-volumetric attacks: - 384 packet-intensive attacks - 329 bit-intensive attacks - 189 request-intensive attacks - Average attack rates reached 3 billion packets per second, 4 Tbps, and 54 million requests per second. - Maximum observed rates reached 9 Bpps, 24 Tbps, and 205 million requests per second. ## Hyper-Volumetric Attacks Reach New Records - Hyper-volumetric attacks increased 40% in Q4 compared with Q3. - Attack sizes grew more than 700% compared with large attacks in late 2024. - One attack reached 31.4 Tbps and lasted only 35 seconds. - Other record-scale attacks reached 205 million requests per second. - Telecommunications, service providers, and carriers were the primary targets, followed by gaming and generative AI services. - Cloudflare infrastructure itself faced HTTP floods, DNS attacks, and UDP floods. ## Most-Targeted Industries and Locations - Telecommunications, service providers, and carriers became the most-attacked industry, replacing Information Technology & Services. - Gambling and casinos ranked third, while gaming ranked fourth. - Computer software and business services climbed significantly in the top-ten rankings. - China, Germany, Brazil, and the United States remained among the most-attacked locations. - Hong Kong rose 12 places to become the second most-attacked location. - The United Kingdom climbed 36 places to rank sixth. Cloudflare’s data shows that organizations should prepare for attacks that combine enormous volume with rapidly changing techniques. Automated, network-scale defenses capable of identifying and adapting to large botnets are becoming a necessity rather than an optional protection.

Read original(opens in new tab)