Techlist.io - Korean Tech Blog Curator

figma2 min readCurated summary

Inside Figma: tips from the team that builds Figma | Figma Blog

Figma’s team uses the product for far more than interface design, including brainstorming, branding, team activities, and remote research. The post presents practical methods and templates that make creative collaboration more approachable, regardless of design experience. Its central recommendation is to use Figma as an open, flexible space for exploration rather than focusing immediately on finding the “right” answer. ## Figma for More Than Product Design - Figmates across engineering, product, design, and research shared favorite workflows in a livestream. - Topics included: - Keyboard shortcuts and speed tips - Brainstorming exercises - Templates for different skill levels - Blend modes and easing curves - The team also uses Figma for team-building and remote user research. ## Brainstorming with Mind Maps Brand designer Remilla Ty recommends a “mind map sketchbook” for gathering inspiration and exploring creative directions. - **Prompt:** The project name or a statement describing the work. - **Image:** Visual references and sources of inspiration. - **Keywords:** Words connected to the images and the original prompt. - The framework helps teams identify themes and turn them into creative concepts. - Figma used this approach while developing branding directions for Figma Community. - The goal is open exploration rather than judging ideas as right or wrong. ## Lists and Mood Boards A second brainstorming method combines structured prompts with visual references. - **Prompt:** A literal description of the project. - **Beyond:** A more abstract interpretation of the project. - **Look and feel:** The emotions the finished work should create. - Teams extract recurring themes and keywords from the lists. - They then build mood boards using images and color palettes. - The resulting visual direction can be distilled into a concept statement and eventually a brand identity. - Figma used this process for its remote Maker Week branding. ## Creating a More Comfortable Research Environment Researcher Nannearl Brown describes using Figma to make participant activities feel less intimidating. - The research team prioritizes helping participants feel comfortable sharing information. - Participants may receive Figma trading cards as pre-work. - The cards encourage them to introduce themselves and prepare for the session. - Questions cover both Figma-related topics—such as favorite features and use cases—and personal information. The practical takeaway is to use Figma’s collaborative canvas not only for polished design work, but also for low-pressure ideation, visual thinking, and human-centered research.

Read original(opens in new tab)
datadog1 min readCurated summary

How we minimized the overhead of Kubernetes in our job system | Datadog

The supplied content does not include the blog post itself; it contains Datadog’s navigation menu and a link titled “Moving a Job System to Kubernetes.” As a result, there is not enough article text to accurately summarize its arguments, implementation details, or conclusions. ## Available Information - The linked post appears to concern migrating a job-processing system to Kubernetes. - The surrounding page lists Datadog products for: - Infrastructure and Kubernetes monitoring - Application performance monitoring - Logs, databases, and jobs - Security and software delivery - No technical discussion, architecture description, challenges, or results from the post is included. Please provide the article’s body or a working page extract for a detailed summary.

Read original(opens in new tab)
datadog3 min readCurated summary

How we minimized the overhead of Kubernetes in our job system

Kubernetes can improve machine management and scalability, but its scheduling and runtime overhead can significantly reduce job throughput if configured poorly. Datadog found its Kubernetes-based job system used more CPU and completed jobs 40–50% more slowly than the previous VM-based system. By designing a controlled experiment, choosing better metrics, and tuning pod resource requests, the team recovered performance to roughly VM parity while investigating the overhead of running one parent process per pod. ## Designing a Comparable Experiment - The initial comparison was difficult because the Kubernetes and VM deployments differed in: - Number of nodes - Number of worker-parent clusters - Workload and enqueue rate - A controlled experiment was created using: - Identical `c5.2xlarge` machines - The same kernel version, `3.13.0-141` - Both systems repeatedly running a simple Python job - Each Kubernetes pod contained one parent process and its worker processes, making pod count equivalent to parent-process count per node. - The older kernel did not include CPU mitigations, avoiding that variable in the comparison. ## Choosing Useful Performance Metrics ### Measuring node effort - Load average initially appeared useful for measuring machine utilization. - Kubernetes background processes—such as cluster polling and pod-state checks—artificially increased load average. - Load average counts runnable processes rather than the amount of CPU time they actually consume. - The team therefore used CPU idle time instead: - It measures unused CPU capacity. - It reflects actual CPU work rather than the number of active processes. ### Measuring system performance - The job system optimized for throughput rather than latency. - Throughput was measured by the number of jobs completed within 30 seconds. - Latency remained useful for detecting queueing problems, but throughput was the primary success metric. ## Tuning Kubernetes Resource Requests - The main performance gains came from improving pod scheduling. - The target was six pods per `c5.2xlarge` node. - Initially, each pod requested: - One full CPU core - More memory than necessary - Since the node had eight cores and approximately 1.5 GiB of memory consumed by Kubernetes and system services, only four pods could be scheduled. - Requests were reduced to: - `100m` CPU, or 100 millicores - `500 MB` memory - CPU tuning generally enabled six pods per node, although some nodes still scheduled only five. - Further memory reduction was needed because system daemons consumed enough memory to prevent six pods from fitting on some nodes. - Resource requests affect scheduling minimums, while limits constrain containers after they start. - These request changes did not slow jobs because the pods still received sufficient resources to operate. ## One Parent Process per Pod - The team considered placing multiple parent processes in each pod to reduce potential pod overhead. - One parent plus its workers was a natural application unit and simplified orchestration. - The decision depended on how much overhead each pod introduced: - High overhead would favor fewer, larger pods. - Low overhead would favor one parent per pod for simpler management. - Using `pstree`, the team identified six job-system instances per node and traced their process trees through components such as: - `containerd-shim` - `tini` - The application process - They estimated that each pod included overhead associated with three containers, particularly `containerd-shim`. - CPU overhead was then investigated using `perf sched`. The practical lesson is to compare equivalent workloads, measure actual CPU consumption rather than relying blindly on load average, and tune Kubernetes requests for the desired packing density. Resource requests should be large enough for reliable operation but not so large that they unnecessarily prevent pods from being scheduled together.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering spotlight: Maël Nison | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The announcement positions Datadog as a broad observability platform spanning infrastructure, applications, logs, security, digital experience, software delivery, service management, and AI. The provided content does not include Gartner’s detailed evaluation or the blog post’s supporting arguments. ## Recognition and Platform Scope - Datadog highlights its leadership placement in Gartner’s observability-platform research. - Its platform covers: - Infrastructure and container monitoring - Application performance monitoring and profiling - Database, data-stream, and jobs monitoring - Log management and observability pipelines - Cloud, application, workload, and code security - Browser and mobile real user monitoring - Synthetic monitoring, session replay, and error tracking - CI visibility, testing, code coverage, and feature flags - Incident response, service catalogs, SLOs, and workflow automation ## AI and Automation - Datadog presents AI as an integrated part of its platform through: - Bits AI agents and investigation tools - AI integrations and agent observability - GPU monitoring - MCP Server and agent-building capabilities - AI-assisted security and developer workflows - Additional automation features include Watchdog, fleet automation, workflow automation, and incident-management tools. ## Overall Positioning - The product catalog emphasizes a unified approach to monitoring technology environments rather than separate tools for infrastructure, applications, security, and user experience. - The platform also includes dashboards, alerts, notebooks, governance controls, access management, and mobile access. The announcement’s central message is that Datadog combines extensive observability coverage with security, delivery, service-management, and AI capabilities. Readers seeking the actual Gartner assessment should consult the linked Magic Quadrant resource, since the supplied text contains only the announcement and navigation information.

Read original(opens in new tab)
datadog3 min readCurated summary

Engineering spotlight: Maël Nison

Maël Nison’s journey from learning DarkBASIC on La Réunion and in Toulouse to becoming Yarn’s principal maintainer illustrates how curiosity, open source, and a focus on solving practical problems can shape a career. His early experimentation with games, websites, forums, and content-management systems developed into a lasting interest in improving developer workflows. That path eventually led through EPITECH, startups, Facebook, and Datadog, while giving him broad experience across both software and community leadership. ## Early Programming on La Réunion and in Toulouse - Maël grew up on the remote Indian Ocean island of La Réunion, where he had little access to computers. - After moving to Toulouse, he discovered a school programming club and began creating games with DarkBASIC. - DarkBASIC simplified 2D and 3D Windows game development through built-in libraries, tutorials, and DirectX support. - Seeing code immediately produce something on screen made programming feel logical and compelling to him. - By high school, he was building PHP websites, working with SQL, and experimenting with multiple languages and platforms. ## Discovering Open Source and Workflow Automation - In the early 2000s, distributing software was much harder because platforms such as GitHub did not yet exist. - Maël shared source archives through online forums, reflecting the informal nature of early open source communities. - His interest in forum software led to work on content-management systems. - He focused on reducing repetitive administrative workflows, such as allowing users to edit content directly instead of navigating through multiple administration pages. - This pattern—identifying a problem, building a solution, and sharing it with others—became a central theme in his career. ## Education at EPITECH - Maël attended EPITECH in Paris, an institution centered on practical technical education and self-directed learning. - The school emphasized peer assessment and hands-on projects rather than traditional, theory-heavy instruction. - He also spent a year abroad in Québec. - During his final year, he combined his studies with his first full-time job, gaining professional experience before graduation. ## Joining Facebook and Yarn - In 2017, after several years in startups, Maël moved from France to London seeking opportunities at larger organizations. - He joined Facebook without specifically intending to work on a package manager. - Facebook’s onboarding “boot camp” identified his skills and connected him with the emerging Yarn project. - He welcomed the opportunity to work on open source during his regular working hours. - What began as a few pull requests became a multi-year role as a major maintainer and leader of the project. ## Yarn’s Technical and Community Evolution - Yarn was rewritten in TypeScript and re-architected into a more modular system. - It evolved from an internal Facebook tool into a genuinely community-driven open source project. - Maël’s responsibilities expanded far beyond coding: - Product management and roadmap planning - Team leadership and infrastructure - Customer support and community work - Web design, evangelism, and outreach - Defining the project’s broader vision - Although he left Facebook for Datadog in 2019, he continued leading Yarn while taking on new challenges at Datadog. Maël’s experience suggests that careers can grow from small, self-directed experiments into major technical leadership opportunities. Developers can follow a similar path by solving concrete problems, sharing their work openly, and being willing to take on the technical, organizational, and community responsibilities that accompany successful projects.

Read original(opens in new tab)
datadog3 min readCurated summary

PHP 8: Observability baked right in

PHP’s observability mechanisms failed to keep pace with Zend Engine improvements in PHP 7 and PHP 8, especially the introduction of JIT. Existing hooks imposed significant runtime costs, created compatibility and stability problems, and limited tracers such as Datadog’s ability to evolve. PHP 8 addressed these issues by introducing a new observer API designed specifically for modern, lower-overhead runtime instrumentation. ## Observability Before PHP 8 ### The `zend_execute_ex` VM Hook - Extensions could override `zend_execute_ex` to intercept every PHP-defined function and method call. - This moved PHP calls onto the native C stack, whose limited size (`ulimit -s`) could cause stack overflows and process crashes. - Every userland call was intercepted, even when an extension only needed to observe a subset, adding overhead to call-heavy applications. - The compiler could no longer use optimized distinctions between userland and internal calls, such as `DO_UCALL` and `DO_ICALL`. - Extensions had to manually forward the hook to other extensions, creating “noisy neighbor” problems, unexpected behavior, and possible crashes. - The hook was incompatible with PHP 8’s JIT compiler. ### Custom Opcode Handlers - Extensions could replace handlers for function-call opcodes, avoiding the native-stack problem associated with `zend_execute_ex`. - These handlers still required careful forwarding to neighboring extensions, which was historically unreliable. - Handlers could mutate VM state—for example, preventing the original opcode from running—making reliable cooperation between multiple extensions impossible in some cases. - Generators could not be fully instrumented through custom opcode handlers. - Like `zend_execute_ex`, custom opcode handlers were incompatible with the PHP 8 JIT. ### Zend Extension Hooks - Zend Extensions had privileged access to engine-level function-call begin and end handlers. - This approach caused the compiler to emit `EXT_FCALL_BEGIN` and `EXT_FCALL_END` around every function call. - The additional opcodes introduced too much overhead for production-grade tracing. ### AST Injection Experiments - Researchers explored injecting observability nodes into the abstract syntax tree during compilation. - These nodes could invoke tracing functions before and after calls. - However, injecting instrumentation around every function call was expected to have overhead comparable to Zend Extension hooks. - No production-ready tracers using this approach were known at the time. ## The Need for a New Observer API - Existing hooks forced observability tools to interfere deeply with VM execution or compiler output. - Their limitations included excessive overhead, stack-safety risks, incomplete generator support, extension conflicts, and JIT incompatibility. - These constraints prevented tools such as the Datadog PHP tracer from taking full advantage of PHP 8. - In response, the authors and the PHP internals community developed and shipped the observer API in PHP 8, providing a foundation for more modern and efficient tracing, profiling, and debugging. PHP 8’s observer API was necessary because older instrumentation techniques were either unsafe, too slow for production, difficult to compose, or incompatible with the JIT. A runtime-level observability mechanism designed alongside the engine is a more sustainable approach than modifying VM hooks, opcodes, or compiled syntax from extensions.

Read original(opens in new tab)
datadog1 min readCurated summary

PHP 8: Observability baked right in | Datadog

Datadog announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The supplied content contains the announcement headline and Datadog’s product navigation, but not the report’s evaluation details, methodology, or supporting arguments. ## Gartner Recognition - Datadog is positioned as a Leader in Gartner’s Magic Quadrant for Observability Platforms. - The linked resource appears to provide the full Gartner report or announcement. - No specific Gartner strengths, cautions, rankings, or comparison with other vendors are included in the provided text. ## Datadog’s Observability Portfolio The page navigation highlights Datadog’s broad platform, including: - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless, and cloud costs - APM, service monitoring, profiling, and dynamic instrumentation - Database, data-stream, job, and quality monitoring - Log management, observability pipelines, and sensitive-data scanning - Real-user monitoring, session replay, synthetic monitoring, and error tracking - CI visibility, testing, code coverage, and software delivery tools - Incident response, service catalogs, SLOs, workflow automation, and event management - AI capabilities such as agent observability, Bits AI, GPU monitoring, and MCP integrations Overall, the material presents Datadog as a broad, integrated observability platform, but the actual Gartner analysis is not included. For a detailed assessment, consult the linked Gartner report directly.

Read original(opens in new tab)
figma3 min readCurated summary

Inside Figma: a case study on strict null checks | Figma Blog

Figma enabled TypeScript’s `strictNullChecks` incrementally to eliminate null-related bugs without halting product development. The migration addressed more than 4,000 compiler errors across roughly 1,162 files and helped prevent a class of production incidents. Figma concluded that a progressive, allowlist-based rollout was more practical than either stopping all development or fixing errors indefinitely without enforcing the setting. ## What Strict Null Checks Provide - Without `strictNullChecks`, ordinary types such as `Vector` may also contain `null`, making unsafe property access possible. - With the option enabled: - Non-nullable types cannot be assigned `null`. - Nullable values must be explicitly declared, such as `Vector | null`. - TypeScript uses control-flow analysis to narrow types after null checks. - This prevents errors such as accessing `.name` on `undefined`. - The type system also documents important assumptions, such as whether data has been loaded, making code easier to maintain. - Figma’s historical incident data showed that strict null checks could have caught several high-severity production issues before release. ## Why the Migration Was Difficult - Figma adopted TypeScript before strict null checks existed and had accumulated code that did not satisfy the newer rules. - Enabling the option immediately produced more than 4,000 errors across approximately 1,162 frontend TypeScript files. - Fixing one error could expose additional errors. - Meanwhile, the codebase continued growing—from 376,000 to 464,000 lines during the migration—making a strategy that allowed progress to reverse particularly risky. ## Alternatives Figma Rejected ### Stop-the-World Migration - All engineers could have paused product work to fix the type errors. - Figma rejected this because: - The work was only partly parallelizable. - Product development was strategically important. - Coordinating a company-wide effort becomes increasingly difficult as organizations grow. ### Whack-a-Mole Error Fixing - Teams could fix strict-null errors over time while enforcing the checks only in CI. - This minimizes disruption but permits new code to introduce additional errors. - The approach is viable only if errors are fixed faster than they are added. - Figma found a continually moving progress target unattractive, especially given rapid codebase growth. ## Progressive Allowlisting - Figma chose to strict-null-check one file at a time by adding successfully migrated files to an allowlist. - The approach was inspired by the VS Code team’s strict-null migration. - The build compiles the codebase twice: - Once with strict null checks disabled for the general codebase. - Once with strict null checks enabled for files on the allowlist. - This allowed teams to migrate files incrementally while continuing normal development elsewhere. Figma’s experience recommends progressive enforcement for large, actively developed TypeScript codebases: isolate compliant files, enforce the stricter rules there, and expand the allowlist until the entire codebase is migrated.

Read original(opens in new tab)
figma2 min readCurated summary

Redesigning Dropbox’s ways of working | Figma Blog

Dropbox’s shift to a virtual-first model required more than moving meetings online; it meant redesigning how teams collaborate and maintain creative energy. Having already adopted Figma, the design team expanded its use from design production to brainstorming, reviews, handoffs, and cross-functional communication. Figma ultimately became a shared visual workspace for teams across Dropbox, helping recreate some of the spontaneity and connection of the physical office. ## The Challenge of Virtual-First Work - Dropbox’s culture and workspaces had been designed around in-person collaboration across twelve global offices. - Design teams relied on physical workshops, post-its, whiteboards, and informal interactions to generate ideas. - Remote work made creative flow and team connection less automatic, requiring more intentional collaboration practices. ## Redesigning Collaboration with Figma - Dropbox’s design organization had migrated from Sketch to Figma in 2018 and already used it for: - Brainstorming - Wireframing - Prototyping - Commenting - During the pandemic, designers increased their reliance on Figma to reduce communication gaps. - Observation mode replaced some in-person design reviews, allowing participants to follow a presenter’s work directly. - Designers shared files for rapid iteration instead of pairing at a desk or brainstorming in a physical room. - Design sprints brought engineers and product managers directly into Figma, making participation easier for non-designers. - Teams also created asynchronous workflows, such as using checkmarks in an icon template to signal that assets were ready for inclusion in the master library. - These file-based processes reduced unnecessary communication and improved the efficiency of recurring library releases. ## Expanding Figma Across Dropbox - The design team became an example of how remote collaboration could work for the broader organization. - Cheechee Lin led workshops teaching cross-functional teams how to navigate Figma and use features such as commenting, exporting, viewing, and following. - More than 200 people from marketing, research, data science, and other teams attended the workshops. - Figma provided a shared visual space that compensated for the lack of physical whiteboards. - Teams beyond design began using Figma for test plans, research presentations, survey data, qualitative insights, and collaborative presentations. Dropbox’s experience suggests that successful virtual-first work depends on deliberately rebuilding collaborative habits, not simply adopting video calls. Making a shared, accessible workspace central to communication can help distributed teams preserve creativity, speed, and cross-functional involvement.

Read original(opens in new tab)
figma2 min readCurated summary

Meet us in the browser | Figma Blog

Figma’s five-year journey demonstrates how browser-based software can transform design from a siloed, offline activity into an open, collaborative practice. Although this shift initially threatened designers’ sense of control and identity, the browser’s multiplayer, accessible nature encourages shared ownership and transparency. Field argues that this ethos should extend beyond design to all software and society. ## The Browser as a Radical Bet - Figma launched by betting on the browser rather than traditional desktop software. - The founders were drawn to browser-native values: - Collaboration - Transparency - Accessibility - Figma offered practical benefits such as cross-platform support, multiplayer editing, and a single source of truth for files. ## Why the Change Felt Threatening - Moving design from an offline, single-player experience into a shared browser environment challenged established workflows. - Designers worried that transparency could lead to: - Over-demystification of the creative process - Increased micromanagement - Tighter deadlines - Loss of control over their work - Broader access also raised an identity question: if anyone can design, what does it mean to be a professional designer? ## From “My Ideas” to “Our Ideas” - Collaborative digital spaces encourage teams to treat work as shared rather than individually owned. - This requires trust, transparency, and a willingness to let others modify or remix personal creative work. - Figma’s multiplayer environment makes collaboration part of the process, especially when people might otherwise hide unfinished or blocked work. ## Digital Spaces and Broader Access - The browser represents a larger transition from physical spaces to digital ones, accelerated by COVID-19. - Unlike physical spaces, digital spaces can be open by default and less hierarchical. - Browser-based tools reduce dependence on expensive hardware, specialized software, and restrictive titles or credentials. - Figma aims to make design accessible throughout the entire process, from ideation through production. ## A Broader Software Philosophy - Field sees the browser as more than a productivity improvement; it is a medium for cultural change. - Software makers have a responsibility to make platforms usable by more people, not fewer. - The ideal is to create environments where people can work together, embrace creative messiness, learn from failure, and leave ego aside. The article recommends building software around openness, accessibility, and collaboration. Figma’s experience suggests that browser-native tools can help replace siloed expertise with shared participation.

Read original(opens in new tab)
figma2 min readCurated summary

Adding it all up: The math behind designing your career | Figma | Figma Blog

A career is less a predetermined trajectory than a collection of experiments, interests, and opportunities whose meaning becomes clearer in hindsight. Kylie Poppen argues that because work occupies roughly 90,000 hours of our lives, we should actively examine the patterns shaping our choices rather than assume a single “destiny.” Mathematical metaphors—scatterplots and Venn diagrams—offer tools for making more intentional career decisions. ## Careers as Scatterplots - People often describe careers as trend lines, connecting past experiences into a coherent narrative. - In reality, careers are lived as scatterplots: a series of moves between interests, jobs, and opportunities. - The next point cannot reliably be predicted from the current one; clarity usually comes only afterward. - Regularly reflecting on experiences can help reveal a more accurate career trajectory over time. - Work deserves careful consideration because it consumes about one-third of waking life and affects financial stability, purpose, frustration, pride, and identity. ## Finding Meaning Through Overlap - Venn diagrams illustrate how combining different interests can reveal unexpected possibilities and innovations. - Career choices do not need to force a decision between passion and practicality; meaningful work can emerge where hobbies, skills, and interests intersect. - Poppen’s childhood interest in designing posters with CorelDRAW, combined with her passions for storytelling, technology, and people, eventually pointed toward product design. - Her path was not direct: she explored recruiting, marketing, engineering, and project management before gaining the confidence to pursue design. - Mentorship and constructive feedback helped her see design as a craft developed through practice rather than an innate destiny. - Mapping personal interests can uncover overlooked career options and provide confidence to pursue a desired direction. A practical approach is to treat your career as an evolving set of data points, then look for recurring patterns and intersections among your interests, abilities, and experiences. This can lead to more fulfilling choices without requiring certainty about the final destination.

Read original(opens in new tab)
figma2 min readCurated summary

Behind the feature: the making of the new Auto Layout | Figma Blog

Figma’s new Auto Layout evolved from a long-standing idea into a more flexible system for creating responsive designs. The team sought to combine flexbox-inspired power with an approachable interface, defining Auto Layout as “a thoughtful subset of flexbox.” The result was a feature designed to reduce manual resizing while preserving usability for designers who may not know CSS deeply. ## The Need for Automatic Resizing - Before Auto Layout, designers had to manually resize buttons, reposition neighboring elements, and adjust containers whenever content changed. - This repetitive work made responsive design systems difficult to build and maintain. - Figma had considered automatic layout since its early designs, but the concept remained unimplemented until Maker Week in 2018. - During Maker Week, product director Sho reimagined the feature in prototypes, eventually leading to a dedicated team focused on making Auto Layout real. ## Designing a Thoughtful Subset of Flexbox - The team drew inspiration from CSS flexbox because it could help align design workflows with implementation in code. - They intentionally avoided reproducing all of flexbox’s complexity. - Their guiding principle was to create “a thoughtful subset of flexbox” that remained easy to learn and use. - Core capabilities included: - Enabling vertical or horizontal layout on a frame. - Setting spacing between items. - Automatically sizing frames to hug their contents along the main axis. - Allowing fixed or content-based sizing on the counter axis. - Giving individual components independent alignment within their container. ## Prototyping and Team Alignment - The designer, Marcin, built an early prototype entirely in HTML. - The prototype helped the team experience the feature early and resolve interaction details before implementation. - It informed decisions about frame-handle decoration, dragging behavior, and visual outlines. - This shared prototype helped product, design, and engineering align around both the feature’s behavior and its usability. Figma’s approach shows that a successful layout tool does not need to duplicate the full complexity of web technologies. By selecting the most useful flexbox concepts and presenting them through an intuitive editor, Auto Layout made responsive design more practical while keeping the experience accessible.

Read original(opens in new tab)
airbnb3 min readCurated summary

Taming Service-Oriented Architecture Using A Data-Oriented Service Mesh

Airbnb’s Viaduct rethinks the service mesh as a data-oriented layer rather than a network for routing procedural service calls. Built on GraphQL, it presents a unified data graph that hides microservice dependencies from consumers and improves modularity in large SOAs. The central schema can also coordinate service APIs, database models, and serverless data transformations, making system-wide changes more agile. ## The Problem with Large SOAs - Modern organizations may operate thousands of microservices connected through highly tangled dependency graphs. - These graphs resemble “spaghetti code” at the service level: - Changes become difficult to plan. - Teams must coordinate across many service boundaries. - Consumers often depend directly on multiple underlying services. - Airbnb argues that microservice architectures need stronger organizing principles and technical mechanisms for enforcing modularity. ## From Procedure-Oriented to Data-Oriented Design - Traditional procedural design groups procedures into modules with public APIs and hidden implementation details. - Data-oriented design instead organizes software around encapsulated data objects and the methods that operate on them. - Microservices have largely returned SOA to a procedural model: - Each service exposes collections of remote procedural endpoints. - Consumers must know which services provide the data they need. - Viaduct applies data-oriented principles to the service mesh itself. ## Viaduct’s GraphQL Data Mesh - Viaduct defines the mesh through a GraphQL schema containing: - Types and interfaces representing managed data. - Queries and subscriptions for reading data. - Mutations for updating data. - The schema forms a single graph spanning data owned by many microservices. - A consumer can navigate related data through one query, such as: - `productById { manufacturer }` - `productById { reviews }` - `productById { reviews { author } }` - Viaduct determines which services provide each requested field. - This hides service dependencies from consumers and prevents every client from building its own cross-service orchestration logic. ## The Central Schema - Unlike distributed GraphQL approaches that split schemas across modules or federated services, Viaduct treats the schema as one central artifact. - Airbnb uses schema-management primitives to let multiple teams collaborate while preserving a unified model. - Portions of the central schema can define individual microservice APIs. - Airbnb ultimately aims to use the same schema to define database structures. - This could improve “data agility”: - Database changes would no longer need manual translation through several API layers. - A single schema update could propagate changes from storage through services to clients. - Cross-team coordination and delivery times could be reduced. ## Serverless Derived Fields - Many SOAs contain stateless services that transform backend data for particular clients or presentation layers. - Viaduct supports derived fields computed by serverless cloud functions. - These functions operate on the graph without needing direct knowledge of the underlying microservices. - Moving transformation logic into stateless containers can: - Reduce the number of services. - Lower operational overhead. - Keep the core service graph simpler. ## Implementation and Operational Features - Viaduct is built on `graphql-java`. - It supports fine-grained field selection through GraphQL selection sets. - It uses data-loading techniques and an intra-request cache. - Reliability features include short-circuiting and soft dependencies. - Field-level observability shows which services consume particular data. - Its GraphQL interface enables use of established open-source tooling and interactive development tools. Viaduct’s practical recommendation is to place a unified data schema at the center of the architecture, allowing the mesh—not individual consumers—to manage service composition. This can make large SOAs more modular, easier to evolve, and better suited to serverless execution.

Read original(opens in new tab)
datadog3 min readCurated summary

Introducing Glommio, a thread-per-core crate for Rust and Linux

Thread-per-core architecture can significantly improve performance and reduce cloud costs by avoiding lock contention and expensive context switches. However, adopting it directly can reduce developer productivity because it requires new programming patterns and careful data ownership. Datadog developed Glommio, a Rust framework intended to make thread-per-core applications easier to build and maintain. ## Why Traditional Threading Has Limits - Applications commonly use multiple threads to perform independent tasks in parallel. - Shared data requires locks, which introduce contention and waiting. - Thread context switches can cost around five microseconds—potentially more than modern storage I/O operations using technologies such as `io_uring`. - Asynchronous programming reduces blocking, but many runtimes still rely on thread pools or separate worker threads for operations such as file I/O. ## How Thread-per-Core Works - Each CPU core runs a single application thread, often pinned to that core. - Because the operating system does not move the thread between cores, ordinary thread context switches are eliminated. - Hardware interrupts and auxiliary tasks can still interrupt execution. - For maximum performance, operators may reserve certain CPUs for interrupts and system services rather than application work. ## Sharding Data Across Cores - Thread-per-core applications depend on sharding: each thread owns a distinct subset of the data or requests. - Examples include assigning Kafka partitions or database key ranges to individual threads. - Requests assigned to one thread execute there to completion unless the code explicitly yields. - This ownership model prevents multiple threads from handling the same request or data simultaneously. ## Eliminating Locks - Since one thread processes a shard at a time, operations on that shard are naturally serialized. - A conventional threaded cache requires locks because multiple threads may update the same data concurrently. - Sharding reduces contention by dividing a large cache into smaller sections, but locks may still be needed if the operating system switches between threads. - With thread-per-core, updates to keys in the same shard occur sequentially, so an update can complete without acquiring a lock. ## Glommio and Existing Precedents - Thread-per-core is not a new concept; the author previously worked with Seastar, a C++ framework used by ScyllaDB. - Datadog’s Glommio brings the model to Rust while aiming to make its programming challenges more manageable. - The framework is motivated by the need to preserve developer productivity while achieving the efficiency gains of thread-per-core systems. Thread-per-core is most suitable for highly parallel, high-throughput workloads with naturally shardable data. Its performance benefits depend on disciplined data ownership and cooperative execution, while frameworks such as Glommio can reduce the complexity of adopting the model.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Glommio, a thread-per-core crate for Rust and Linux | Datadog

Datadog has been recognized as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The announcement positions Datadog as a broad observability provider spanning infrastructure, applications, logs, security, digital experience, software delivery, service management, and AI. The supplied content contains mostly site navigation rather than the article’s supporting details or Gartner’s evaluation rationale. ## Datadog’s Observability Scope - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless systems, cloud costs, GPUs, and storage. - Application performance monitoring, service monitoring, profiling, dynamic instrumentation, and agent observability. - Log management, sensitive-data scanning, audit trails, and observability pipelines. - Database, data-streams, data-quality, and jobs monitoring. ## Broader Platform Capabilities - Security features including cloud security, SIEM, workload protection, code security, vulnerability management, and compliance. - Digital-experience tools such as real-user monitoring, session replay, synthetic monitoring, error tracking, and product analytics. - Software-delivery capabilities covering CI visibility, test optimization, continuous testing, code coverage, and feature flags. - Service-management tools for incidents, events, SLOs, workflows, case management, and software catalogs. - AI offerings including Bits AI agents, investigation tools, agent observability, GPU monitoring, and MCP integrations. Overall, the announcement emphasizes Datadog’s unified and expansive observability platform. A complete assessment of Gartner’s specific strengths, cautions, and evaluation criteria would require the full blog post or linked Gartner report.

Read original(opens in new tab)