Techlist.io - Korean Tech Blog Curator

figma2 min readCurated summary

Unlocking the Power of Code Connect | Figma Blog

Code Connect helps design systems bridge the gap between design and engineering by making implementation guidance and production code available directly in Figma’s Dev Mode. Leaders from Bumble, GitHub, and HP argue that successful adoption depends on creating shared terminology, meeting developers in their existing workflows, and introducing the tool incrementally. The broader conclusion is that Code Connect is a practical first step toward tighter, more consistent design-to-code collaboration. ## Establish a shared language - Designers and developers often use different naming conventions, component properties, and expectations. - Design systems can serve as a “third language” by defining shared terms, patterns, and components. - Documentation and design decisions must be easy to discover and implement, not merely documented somewhere. ## Bring design guidance into developer workflows - Teams often struggle to identify a single source of truth, leading to inconsistent implementations and duplicated code. - Designers may rely on design documentation while engineers consult separate technical sources. - Code Connect surfaces design-system documentation and implementation details directly in Dev Mode, reducing the need for developers to switch to external websites or tools. - This helps teams reuse established patterns instead of repeatedly creating custom solutions. ## Start small to encourage adoption - Design systems cannot assume that people will adopt them simply because they exist; reducing friction is essential. - Code Connect lets developers inspect how a design has been implemented without leaving their existing workflow. - Teams should begin with small, high-impact components such as toggles. - Mapping a few components first makes it easier to learn the process, demonstrate value, and expand gradually. - HP’s Gilson Hoffmeister describes Code Connect as a first step toward broader design-system improvements. ## Value both disciplines’ expertise - Effective design-to-code collaboration requires recognizing that designers and developers contribute different forms of expertise. - Code Connect is intended to connect those perspectives by delivering design-system code into Dev Mode. - Better alignment can improve consistency, speed, and the quality of handoff without requiring either discipline to abandon its established practices. Teams should begin with a small set of frequently used components, establish shared terminology and ownership, and use Code Connect to make authoritative implementation guidance available where developers already work.

Read original(opens in new tab)
figma2 min readCurated summary

Speeding Up File Load Times, One Page At A Time | Figma Blog

Figma improved file-loading performance by dynamically loading only the page a user opens instead of the entire file. This approach reflects user-perceived complexity, reduces memory usage, and avoids making small pages wait for unrelated content. For the slowest 5% of loads, the change reduced load times by 33%. ## Loading Content According to User Needs - Figma files can contain dozens of pages, hundreds of frames, components, styles, variables, and prototype screens. - Usage data showed that users often treat a file as an entire project but typically visit only a subset of its pages in one session. - Loading everything up front made a small page unnecessarily depend on the size of the whole file. - Dynamic loading lets Figma display the selected page first and fetch additional content only when needed. ## Cross-Page Read Dependencies - A Figma file is modeled as a tree of nodes, with nodes representing interactable layers and their properties. - Nodes can reference content located on other pages, creating cross-page dependencies. - An instance points to its backing component, which may be on another page; the component must be loaded before the instance can render correctly. - Styles and variables also create dependencies: - A fill style requires loading its corresponding style node. - A variable-based font size requires the variable node so the client can resolve the raw value. ## QueryGraph and Earlier Dynamic Loading - Figma had already developed dynamic loading for view-only files and prototypes. - Its QueryGraph framework stores dependency relationships as an in-memory graph. - The multiplayer system uses this graph to determine which parts of a file should be sent to connected clients. - Previous loading strategies included: - **Page-based canvas loading:** Load the selected page and its required dependencies, then fetch other pages on demand. - **Frame-based prototype loading:** Load the current prototype screen and preload reachable frames within a limited number of transitions. ## Performance Impact - The goal is for load times to trend downward even as files become larger and more feature-rich. - Dynamic loading improves both initial responsiveness and memory consumption. - The largest benefits appear in worst-case loads, with a reported 33% reduction for the slowest 5% of page loads. Figma’s approach demonstrates that large collaborative documents should be loaded according to the user’s immediate context, while dependency tracking ensures referenced content remains available when required.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering VP spotlight: Ivo Dimitrov | Datadog

Datadog announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The supplied content primarily consists of the announcement link and Datadog’s product navigation, so it does not provide Gartner’s evaluation details or the reasoning behind the placement. It does show the breadth of Datadog’s unified observability, security, software delivery, and AI platform. ## Gartner Recognition - Datadog is identified as a “Leader” in Gartner’s Magic Quadrant for Observability Platforms 2026. - The linked resource appears to contain the full announcement and Gartner report, but its substantive text is not included here. ## Broad Observability Platform - **Infrastructure:** Infrastructure, container, Kubernetes, network, serverless, GPU, storage, and cloud cost monitoring. - **Applications and data:** APM, universal service monitoring, profiling, dynamic instrumentation, database monitoring, data streams, jobs, and quality monitoring. - **Logs:** Log Management, Sensitive Data Scanner, Audit Trail, and Observability Pipelines. - **Digital experience:** Browser and mobile RUM, session replay, product analytics, synthetic monitoring, mobile testing, and error tracking. ## Security and Software Delivery - Security offerings include code security, SAST, SCA, cloud security, SIEM, workload protection, vulnerability management, compliance, and application/API protection. - Software delivery tools cover CI visibility, test optimization, continuous testing, code coverage, feature flags, IDE plugins, and internal developer portals. - Service-management capabilities include incident response, event management, SLOs, case management, workflow automation, and service catalogs. ## AI and Platform Capabilities - Datadog highlights AI features such as Bits AI agents, investigation tools, AI integrations, MCP Server, and agent observability. - Platform features include dashboards, alerts, notebooks, Watchdog, access control, governance, fleet automation, and mobile access. Datadog’s positioning is as a consolidated platform spanning telemetry, application and infrastructure monitoring, security, developer workflows, and AI operations. For the specific Gartner assessment, readers would need to consult the linked announcement or report.

Read original(opens in new tab)
datadog3 min readCurated summary

Engineering VP spotlight: Ivo Dimitrov

Ivo Dimitrov’s career evolved from low-level systems programming into engineering leadership focused on large-scale distributed storage. His experience at Microsoft and LinkedIn shaped his approach to building scalable data platforms, while Datadog attracted him with its talented people, modern technology, and culture of experimentation. Today, he leads Datadog’s Distributed Data Systems organization, supporting the company’s metrics, events, query, alerting, and analytics infrastructure. ## From Electrical Engineering to Systems Programming - Dimitrov initially studied electrical engineering and became interested in software while working on digital control systems. - His early work included contributing to a real-time operating system kernel. - He spent roughly a decade as an individual contributor, specializing in: - High-performance systems - Low-level programming - C and C++ - System software ## Transition from Individual Contributor to Manager - At Microsoft, Dimitrov worked on an early version of Azure Blob Storage. - Following a reorganization, he accepted an opportunity to lead his team despite having no prior management experience. - Microsoft supported the transition through: - Leadership mentorship - Formal management training - Guidance on communication, conflict resolution, and interpersonal leadership - He discovered that management allowed him to expand his ownership beyond individual projects and influence broader organizational outcomes. - The role combined his technical background with responsibilities such as cross-functional coordination, team development, and engineering strategy. ## Building Internet-Scale Storage at Microsoft and LinkedIn - At Microsoft, Dimitrov worked on storage systems supporting Hotmail. - After joining LinkedIn in 2014, he adapted to a technology environment centered on open source tools such as MySQL and Java. - He led development of Espresso, LinkedIn’s proprietary key-value storage platform. - The platform matured into a core system supporting approximately 95 percent of LinkedIn’s data sets. - He also helped oversee several other large-scale storage projects: - **Venice**, an open source platform for serving derived data - **Ambry**, an open source blob storage system - **Helix**, an open source cluster manager - These systems supported critical parts of LinkedIn’s internet-scale infrastructure. ## Why Datadog Was Appealing - Dimitrov was drawn to Datadog by three main factors: - Highly capable engineers and leaders - Interesting, modern technology - The opportunity to contribute to a rapidly growing company - Compared with the legacy systems and processes that had accumulated at LinkedIn, Datadog offered less bureaucracy and more freedom to: - Take thoughtful risks - Experiment - Deliver quickly - Fail fast and learn - Iterate and innovate - He was particularly interested in Datadog’s Kubernetes-based Metrics and Events platforms and the challenge of building best-in-class infrastructure during the company’s growth. ## Distributed Data Systems at Datadog - Dimitrov leads the Distributed Data Systems organization, which owns a portfolio of storage and data technologies. - Its responsibilities include: - **Metrics**, supporting metrics and time-series data - **Events**, handling semi-structured data such as logs, profiles, and traces - **Driveline**, a main-memory database optimized for online analytics - The **Cross-Platform Queries** team provides a unified query interface across systems that historically exposed separate, domain-specific APIs. - This reduces the learning curve for engineers and customers. - It abstracts the underlying data stores behind a common API. - The organization also operates Datadog’s Alerts platform, which generates a large share of the queries sent to the Metrics and Events systems. Dimitrov’s experience demonstrates how deep systems expertise can translate into effective engineering leadership. His recommendation by example is to remain technically engaged while expanding one’s scope—from writing individual components to shaping teams, platforms, and long-term engineering direction.

Read original(opens in new tab)
datadog3 min readCurated summary

.NET Continuous Profiler: Memory usage

Datadog’s .NET memory profiler helps identify excessive garbage collection, allocation hotspots, and objects that remain in memory after collection. It combines CLR events, operating-system thread metrics, sampled allocation data, stack traces, and weak handles to provide production-friendly memory insights. The approach favors low overhead, though some capabilities depend on the .NET version. ## Measuring Garbage Collector CPU Impact - The profiler uses CLR events to monitor garbage collection phases. - In server GC mode, the CLR creates two high-priority threads per heap/core to process collections in parallel. - Since .NET 5, these threads are named `.NET Server GC` and `.NET BGC`. - At each profile export, the profiler retrieves these threads’ CPU usage from the operating system. - It records the result as a sample with a native stack containing a `Garbage Collector` frame. - This uses a pull model: the exporter periodically requests the CPU measurement because no suitable event or dedicated profiler thread exists. - Before .NET 5, GC thread CPU usage could not reliably be identified because `GCCreateConcurrentThread` did not include thread IDs. ## Sampling Allocations - Per-allocation callbacks such as `ICorProfilerCallback::ObjectAllocated` provide detailed data but significantly slow allocation fast paths. - `GCSampledObjectAllocation` and `ObjectsAllocatedByClass` reduce some costs but do not provide call stacks for individual allocation sites. - Datadog instead listens to `AllocationTick`, emitted for roughly every 100 KB allocated. - Each event includes: - The object’s `ClassID` and type information. - The allocation address. - The object size and total allocation size since the previous tick. - The allocation kind: SOH (`0`), LOH (`1`), or POH (`2`). - Generic type names are reconstructed through the .NET profiling API. - Because allocation events are synchronous, the current thread is responsible for the allocation; the profiler walks that thread’s stack to capture the allocation call site. - This produces sampled allocation data for each heap category without imposing the cost of observing every allocation. ## Tracking Objects That Survive Garbage Collection - An allocation address alone cannot track an object indefinitely because compacting garbage collections can move objects. - Datadog uses weak handles, created through `GCHandle.Alloc`, which move with objects and do not keep them alive. - The profiler added this functionality through the .NET 7 `ICorProfilerInfo13` API and its `LiveObjectsProvider`. - For every sampled allocation, it creates a weak handle and records the object’s creation time. - After each garbage collection: - Handles for unreachable objects are removed and destroyed. - Handles for surviving objects remain and are included in the next profile. - This lets users inspect representative objects that persist after collection and investigate potential memory leaks. ## Practical Recommendation Use allocated-memory profiles to find endpoints and types responsible for excessive allocation, then examine surviving-object samples for retention or leak investigations. GC CPU data is especially useful for diagnosing applications whose high CPU usage is driven by frequent or expensive garbage collections.

Read original(opens in new tab)
datadog1 min readCurated summary

.NET Continuous Profiler: Memory usage | Datadog

Datadog is presented as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. However, the provided content contains only the announcement headline, link, and website navigation; it does not include the blog post’s analysis, criteria, or supporting details. ## Announcement - Datadog’s headline claim is recognition as a “Leader” in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked page appears to be a Datadog resource or announcement page. ## Available Product Scope The navigation indicates that Datadog’s observability platform spans: - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless systems, and cloud costs - Application performance monitoring, profiling, and dynamic instrumentation - Logs, databases, data pipelines, and data quality - Security monitoring and cloud security - Real user monitoring, synthetic monitoring, session replay, and error tracking - CI/CD visibility, testing, developer portals, and software delivery - Incident management, service catalogs, SLOs, workflow automation, and AI-powered investigation Because the actual article text is missing, no further claims about Gartner’s evaluation or Datadog’s strengths can be reliably summarized.

Read original(opens in new tab)
figma2 min readCurated summary

Keeping It 100(x) With Real-time Data At Scale | Figma Blog

Figma’s LiveGraph powers real-time collaboration by subscribing to GraphQL-like queries and updating clients automatically. Rapid growth—tripled sessions since 2021 and fivefold view-request growth in one year—exposed limits in its single-server, mutation-based architecture. Figma launched “LiveGraph 100x,” a redesign focused on scaling reads and database updates while preserving performance and enabling a safe migration. ## LiveGraph’s Role in Figma - LiveGraph keeps data synchronized across collaborative features such as: - File editing - Comments - FigJam voting - It exposes a web API for subscribing to GraphQL-like queries. - Results are returned as JSON trees based on a schema of entities, relationships, and views. - A custom React Hook automatically re-renders interfaces when subscribed data changes. ## The 100x Scaling Initiative Figma’s growing user base increased both the number and cost of LiveGraph client sessions. At the same time, the underlying database evolved from one PostgreSQL instance into vertically and horizontally sharded infrastructure. The redesigned system needed to: - Preserve or improve service-level objectives for initial loads and updates. - Support more database shards reliably and efficiently. - Scale reads and database-update processing independently. - Allow incremental, transparent migrations without disrupting users. ## Limitations of the Original Architecture Originally, LiveGraph consisted of: - A single LiveGraph server. - An in-memory query cache. - One PostgreSQL instance. - A cache that tailed PostgreSQL’s logical replication stream. PostgreSQL writes row mutations to its write-ahead log, including pre- and post-row images and a monotonically increasing sequence number. LiveGraph used these mutations to update cached query results directly rather than recomputing them. This design worked well at smaller scale because: - All updates came from one primary database. - The replication stream provided a global ordering. - Each row mutation could be applied directly to the relevant cached results. ## Sharding Breaks Global Ordering As the original PostgreSQL instance reached capacity, Figma introduced vertical shards and began moving toward broader horizontal scaling. This invalidated the assumption that all database updates arrive in one globally ordered stream. - Multiple shards can generate updates simultaneously. - Their updates have no guaranteed global order. - LiveGraph therefore needed an architecture that could process distributed database changes while maintaining reliable, timely query updates. The growing load made it necessary to rethink LiveGraph fundamentally rather than continue extending its single-database design.

Read original(opens in new tab)
microsoft2 min readCurated summary

Developing with Accessibility in Mind at Microsoft

Global Accessibility Awareness Day highlights the importance of building inclusive digital products. The post recommends integrating accessibility testing throughout development using Accessibility Insights for Web and Visual Studio’s Integrated Accessibility Checker. Combining automated scans with manual testing helps developers identify both common and deeper accessibility problems. ## FastPass for Rapid Automated Testing - Accessibility Insights for Web uses axe-core to detect common, high-impact accessibility issues. - FastPass can identify problems in under five minutes, often revealing failures within a couple of minutes. - Developers can use it while writing UI code to find and fix issues early. - The tool also includes WCAG 2.2 guidance and testing support in its Assessment feature. ## Visual Studio’s Integrated Accessibility Checker - Available since Visual Studio 2022 version 17.5, the checker scans desktop applications within the IDE. - It detects common accessibility issues and reports them directly in Visual Studio. - The feature is powered by the Axe-Windows engine, also used by Accessibility Insights for Windows. ## Manual Testing with Quick Assess - Automated tools cannot detect every accessibility issue, so manual inspection remains necessary. - Quick Assess provides 10 assisted tests for issues beyond automated detection. - Tests include explanations of why each issue matters, along with remediation resources and examples. - Examples include checking heading levels and reviewing individual instances for easier validation. ## Building Accessibility into Development - Accessibility testing should be part of the product life cycle rather than a final checklist. - Developers can use FastPass’s Tab Stops test to evaluate keyboard navigation and focus order. - Poor focus order can make interfaces difficult to use for people relying on screen readers, magnifiers, or those with reading disorders. - Small, consistent testing practices can significantly improve the experience for users with disabilities. The recommended approach is to start with automated checks, supplement them with Quick Assess and keyboard-based manual testing, and continue improving accessibility throughout development.

Read original(opens in new tab)
microsoft3 min readCurated summary

Copy-on-Write performance and debugging

Dev Drive’s ReFS-based copy-on-write (CoW) linking can significantly improve build performance, though results vary by repository structure. The largest gains occur when builds repeatedly copy assemblies or generate microservice layouts; C++-heavy builds generally benefit less. The post also explains how to inspect CoW links, use performance tools safely, and repair leaked ReFS references. ## Build Performance Results - Testing compared NTFS and Dev Drive on the same Dev Box VM. - Many repositories achieved build-time reductions of 10% or more, with a maximum observed improvement of 43%. - The strongest benefits appeared in: - C# repositories with deep project-to-project dependencies, where MSBuild repeatedly copies assemblies. - Builds that copy many files to construct microservice output layouts. - C++ repositories generally saw smaller improvements because: - MSBuild copies output files less frequently. - MSVC produces fewer, larger files, reducing the impact of lower file-I/O overhead. - Repositories with long chains of large dependent projects benefited less, since serial build stages limited the effect of faster I/O. - Tests used clean source and output directories, separated package restore and tests from build measurements, and ran at least five iterations while excluding the first cold-cache run. - The tests used the `Microsoft.Build.CopyOnWrite` SDK and, where relevant, an updated `Microsoft.Build.Artifacts` SDK. CoW-in-Win32 was not yet available during testing. ## Identifying CoW Links - CoW links, also called block clones, allow multiple files to reference the same physical disk blocks. - `fsutil file queryExtentsAndRefCounts <file>` displays the file’s extents and reference counts. - A reference count such as `Ref: 0x4` indicates that the underlying blocks are shared by four cloned files. - Each cloned file also requires a small amount of metadata storage, typically one additional cluster. ## Using ProcMon on Dev Drive - Dev Drive restricts file-system filter drivers through an allow-list. - To use ProcMon: - Check the current filter list with `fsutil devdrv query`. - Add ProcMon’s current filter driver, such as `ProcMon24`, using `fsutil devdrv setfiltersallowed`. - Dismount the Dev Drive for the change to take effect. - ProcMon’s filter is attached only while ProcMon is running, so it can generally remain on the allow-list. ## Using Microsoft Performance Recorder - Microsoft Performance Recorder requires the `FileInfo` filter driver. - Add `FileInfo` to the Dev Drive filter allow-list and dismount the volume before recording. - Remove `FileInfo` afterward because it remains attached whenever the filter is allowed and can reduce Dev Drive performance. ## Repairing Leaked CoW References - ReFS limits a data block to 8,176 clones. - Excessive or orphaned references can cause errors such as: - `MaxCloneFileLinksExceededException` - `ERROR_BLOCK_TOO_MANY_REFERENCES` (347) - `STATUS_BLOCK_TOO_MANY_REFERENCES` (`0xC000048C`) - The issue is uncommon but can occur after prolonged CoW-heavy builds, particularly with prerelease implementations. - Run `refsutil leak <drive> /s <scratch-file>` from an elevated console to scan and repair dangling references. - Add `/d` to detect leaks without fixing them. - Large volumes may contain billions of leaked references, and the repair process can take considerable time. Dev Drive and CoW linking are most worthwhile for build systems dominated by repeated file copying, especially large C# and microservice-oriented repositories. Teams should also configure diagnostic filter drivers carefully and periodically use `refsutil` if clone-reference errors appear.

Read original(opens in new tab)
figma2 min readCurated summary

Craft and Beauty: The ROI of Marrying Form and Function | Figma Blog

Craft and beauty are presented as business advantages, not merely aesthetic preferences. Leaders from Stripe, Linear, and Figma argue that thoughtful details improve usability, engagement, conversion, and product differentiation. Their examples—including a 20% increase in email conversion and 11.9% higher average revenue for Stripe Checkout users—show that investment in quality can directly support growth. ## Craft Improves Usability and Conversion - Beauty can make products feel easier and more effective, an effect known as the **aesthetic usability effect**. - Clearer language, stronger visual hierarchy, and a smoother journey helped Stripe increase conversion from an email series by **20%**. - Every user touchpoint can either add or subtract value from the overall product experience. - Stripe’s Optimized Checkout Suite, whose quality and detail were key priorities, helped businesses generate **11.9% more revenue on average**. ## Craft as a Mindset, Beauty as the Result - Linear’s Karri Saarinen distinguishes between: - **Craft:** the mindset and care applied while creating something. - **Beauty:** the visible experience and quality that result. - Quality extends beyond appearance. A well-designed object or product should function smoothly, reliably, and pleasantly. - The central question is whether a team is committed to doing excellent work or merely completing tasks as quickly as possible. ## Performance and Quality Are Foundational - Figma views product quality as a hierarchy of needs. - Fundamental qualities such as web performance and a consistent **60 frames-per-second** experience must come first. - These technical foundations enable users to appreciate higher-level design and interaction details. - Craft also involves making products feel intuitive—so well-designed that users feel everything simply works. ## Craft Requires a Company-Wide Culture - Craft is not the responsibility of a single design team; engineering, product, design, and other functions all contribute. - Figma and Linear emphasize that quality should be an inherent cultural value rather than something enforced only through OKRs. - Stripe uses “friction logging” and multidisciplinary “walking the store” exercises, in which teams experience the product end to end as users do. - This approach helps teams identify and remove problems across the full customer journey. The practical recommendation is to treat craft as an organization-wide operating principle. Companies should invest in performance, clarity, usability, and detail at every touchpoint, because form and function together can create both a better experience and measurable business results.

Read original(opens in new tab)
datadog1 min readCurated summary

How we built the Datadog heatmap to visualize distributions over time at arbitrary scale | Datadog

The supplied content does not include the blog post itself; it mainly contains Datadog’s navigation menu and a promotional link announcing its Gartner recognition. The only identifiable article reference is a post about building Datadog’s heatmap for visualizing distributions over time at arbitrary scale, but its body text is missing. ## Available Content - Datadog is promoting its recognition as a Leader in the Gartner® Magic Quadrant™ for Observability Platforms. - The navigation lists Datadog products across: - Infrastructure and application monitoring - Logs, databases, and data observability - Security - Digital experience monitoring - Software delivery - Incident and service management - AI and platform capabilities - The referenced engineering post appears to discuss: - A heatmap visualization - Distributions over time - Scaling to arbitrary data volumes A meaningful technical summary would require the actual article text, which is not present in the supplied content.

Read original(opens in new tab)
datadog3 min readCurated summary

How we built the Datadog heatmap to visualize distributions over time at arbitrary scale

Datadog uses DDSketch-powered distribution metrics and heatmaps to reveal performance patterns that percentile lines can hide. Heatmaps preserve the full shape of latency distributions across hosts and time, making distinct behavioral modes, seasonality, and outliers visible. The visualization is designed to remain scalable and readable even with hundreds of trillions of underlying datapoints. ## Why Unaggregated Distributions Matter - Line graphs reduce billions of events to a single value, such as p50, p99, or max. - Multiple percentile lines provide more context, but the selected percentiles remain arbitrary and can obscure important behavior. - Aggregated percentile changes may suggest that all requests are slowing when only one subset of traffic is changing. - Heatmaps expose separate “modes”—distinct groups of measurements with different behavior. - For example, periodic latency spikes may come from a low-latency benchmarking service rather than from a general degradation in the endpoint. - Filtering out an identified mode can reveal other patterns, such as daily seasonality in the remaining traffic. ## Building Heatmaps with DDSketch - DDSketch sacrifices a small amount of precision to represent extremely large numbers of observations efficiently. - Datadog sends histogram bins and counts to the frontend instead of transmitting every individual datapoint. - Limiting the number of bins keeps the payload size constant as traffic volume grows. - Counts use `float32`, supporting values up to approximately `3 × 10^38` per bin—far beyond practical monitoring volumes. - This allows heatmaps to represent massive datasets, including hundreds of trillions of datapoints. ## Preserving Resolution and Avoiding Aliasing - Heatmap requests contain time buckets, distribution bins, and counts. - Since bucket boundaries are shared across a request, Datadog stores those boundaries only once. - Boundaries must be explicit because distributions may use logarithmic rather than linear scales. - Time buckets need to align with the source data intervals. - Misaligned intervals create aliasing artifacts: for example, grouping 10-second data into 7-second buckets produces repeating count patterns such as `[1, 1, 2, 1, 1, 2, …]`. - Careful discretization preserves the resolution available in the original DDSketch data. ## Designing the Color Scale - The default palette begins with light blue, consistent with other single-series Datadog visualizations. - It transitions toward purple to match Datadog’s visual identity. - The scale avoids lingering on red, which can imply negative alerts, and ends in orange for the hottest values. - Color choices must communicate both the volume and structure of the distribution. ## Maintaining Dynamic Range - A few high-count bins can dominate a linear color scale, leaving most of the heatmap visually indistinguishable. - This is especially problematic for power-law distributions with a dense central mode and a long tail. - A linear scale may clearly show the main mode around 20 ms while hiding a smaller mode near 1 second. - Human brightness perception is nonlinear, approximately following a power law described by Stevens’ law. - Applying nonlinear color interpolation improves the visibility of meaningful differences across both dense regions and long tails. - This helps preserve distribution details that would otherwise be lost when the color range is dominated by outliers or highly concentrated buckets. Datadog’s heatmap approach combines DDSketch compression, aligned high-resolution buckets, and perceptually informed color scaling. For systems where averages or a handful of percentiles conceal important subpopulations, distribution heatmaps provide a more reliable way to investigate performance at scale.

Read original(opens in new tab)
figma3 min readCurated summary

Figma’s journey to TypeScript | Figma Blog

Figma migrated its custom Skew programming language to TypeScript after its original performance advantages became less important than its maintenance and onboarding costs. Advances in mobile WebAssembly support, C++ engine integration, and team growth made the transition practical without sacrificing significant performance. The team completed the migration through an automated, gradual rollout that preserved development velocity and minimized production risk. ## Why Figma Moved Away from Skew - Skew originally helped Figma support prototype viewing across web and mobile. - Its compiler provided optimizations such as: - Constant folding - Devirtualization - Efficient JavaScript integer operations - Fast compile times - Over time, Skew became difficult to scale because: - New engineers struggled to learn it. - It integrated poorly with the broader codebase. - It lacked an external developer ecosystem. - Maintaining its specialized tooling outweighed its performance benefits. - TypeScript offered native package management, static imports, modern language features, extensive tooling, and easier hiring and onboarding. ## Why the Migration Became Possible - Mobile browsers gained broad WebAssembly support by 2018, with reliable performance by 2020. - Figma moved many performance-critical Skew components—especially hot paths such as file loading—to its C++ engine. - These C++ replacements reduced the performance penalty of moving less-critical code to TypeScript. - Larger prototyping and mobile teams provided enough capacity to invest in automated migration tooling. ## Addressing Performance Concerns - In 2020, early benchmarks showed TypeScript could make prototype loading nearly twice as slow in Safari. - Safari was especially important because WebKit was the only browser engine permitted on iOS at the time. - Improved WebAssembly support and the shift of core engine work to C++ made Skew’s compiler optimizations less essential. - Figma gained confidence that TypeScript could provide acceptable performance without recreating Skew’s custom compiler. ## Automated Skew-to-TypeScript Conversion - Manually rewriting the entire codebase would have disrupted development and increased the risk of runtime bugs and regressions. - Figma built a transpiler that converted Skew into TypeScript, extending earlier work by former CTO Evan Wallace. - The migration required care because Skew and TypeScript had different runtime semantics. - For example, TypeScript initializes namespaces and classes only after a module is imported, while Skew made symbols available globally when the codebase loaded. Unexpected import order could therefore introduce runtime failures. ## Three-Phase Rollout ### Phase 1: Write Skew, Build Skew - Figma kept the existing build process. - The new transpiler generated TypeScript from Skew. - Generated TypeScript was checked into GitHub so developers could inspect and prepare for the future codebase. ### Phase 2: Write Skew, Build TypeScript - Once the generated bundle passed unit tests, production traffic began using the TypeScript build. - Developers continued writing Skew. - The transpiler updated the TypeScript source automatically. - The team fixed type errors incrementally; TypeScript could still produce valid bundles despite those errors. ### Phase 3: Write TypeScript, Build TypeScript - After the team adopted the TypeScript build, the generated code became the source of truth. - Figma stopped automatic generation, deleted the Skew source, and required new development to use TypeScript. - The staged process allowed the team to detect and resolve issues such as a Smart Animate regression before completing the cutover. ## Practical Lessons - A custom language can provide valuable early advantages but become a long-term developer-experience liability. - Automated conversion is safer when paired with staged production rollouts and reversible adoption gates. - Controlling the original compiler made it possible to adapt the migration tooling to the codebase’s specific needs. - Figma’s approach demonstrates that large language migrations can preserve delivery speed when technical, performance, and organizational prerequisites are addressed first.

Read original(opens in new tab)
figma3 min readCurated summary

Speeding Up C++ Build Times | Figma Blog

Figma cut C++ build times roughly in half by addressing unnecessary header inclusion rather than relying solely on faster hardware or caching. The team found that compiled bytes were growing much faster than the codebase itself, making transitive header dependencies the main culprit. They combined automated include analysis with CI-based measurement to prevent both unused includes and costly dependency regressions. ## Why Build Times Were Getting Worse - In 2023, Figma’s codebase grew by about 10%, but build times increased by 50%. - C++ builds were a major productivity problem and a top concern in internal developer surveys. - Faster M1 Max machines, Ccache, and remote caching provided only temporary or insufficient improvements. - The team observed that build times were closely related to the amount of code passed to the compiler after preprocessing. ## How C++ Header Inclusion Affects Builds - The preprocessor expands every `#include` into a single large file before compilation. - Transitive dependencies are included as well: - If file C includes B, and B includes A, C receives the contents of both A and B. - As a result, a small source change can cause the compiler to process a very large amount of unrelated code. ## Removing Unnecessary Includes - Figma suspected that many files included headers they did not use directly or relied on headers only for transitive dependencies. - Removing unnecessary includes from the largest files produced: - A 31% reduction in compiled bytes. - A 25% reduction in cold build time. - These results confirmed that compiled byte volume was strongly correlated with build performance. ## DIWYDU: Automating Include Cleanup - Google’s Include What You Use (IWYU) tool was considered but proved difficult to apply retroactively to Figma’s large codebase. - Figma created a less strict alternative called **Don’t Include What You Don’t Use (DIWYDU)**. - DIWYDU: - Uses Python bindings for `libclang`. - Parses source and header files into Clang Abstract Syntax Trees. - Identifies types, functions, and variables directly used by each file. - Flags headers that are included but provide no directly used symbols. - The tool runs on feature branches to prevent unnecessary includes from accumulating. ## DIWYDU’s Limitations - It analyzes Figma-owned files but excludes Standard Template Library headers. - STL headers may define symbols through private internal includes, making direct dependency analysis difficult. - Python’s `libclang` bindings expose less of Clang’s AST than the compiler’s native C++ APIs, sometimes producing `UNEXPOSED_EXPR` nodes. - A future C++ implementation could provide more accurate AST access. - DIWYDU cannot detect cases where an included header is genuinely required but excessively large. - Such regressions may need forward declarations or header decomposition instead. ## Measuring Dependency Growth with `includes.py` - Figma built `includes.py` to measure the transitive bytes associated with each source file. - The tool is written entirely in Python and typically runs in a few seconds without invoking Clang. - It: - Crawls first-party source, header, and generated files. - Counts file sizes. - Builds a dependency graph. - Estimates the total bytes passed to the compiler for each source file. - Standard library includes are treated as zero bytes because Figma mainly accesses them through internal wrapper directories. - CI uses the measurements to compare pull requests and warn authors when changes significantly increase compiled bytes. Figma’s approach demonstrates that controlling header dependencies can deliver larger and more durable gains than simply adding hardware or cache capacity. Teams working on large C++ codebases should automate unused-include checks, measure transitive dependency size in CI, and use forward declarations or smaller headers when necessary.

Read original(opens in new tab)
datadog1 min readCurated summary

How we brought Datadog's data visualization to iOS: A focus on performance | Datadog

The provided content does not include the blog post’s actual article text. It contains Datadog’s navigation menu, product links, and a reference to an engineering post about bringing Datadog data visualization to iOS performance, so its technical argument and conclusion cannot be reliably summarized. ## Available Content - Datadog promotes its recognition as a Leader in the Gartner Magic Quadrant for Observability Platforms. - The navigation lists products across: - Infrastructure and application monitoring - Logs, databases, and data observability - Security - Digital experience monitoring - Software delivery - Service management - AI and platform capabilities - The referenced engineering URL suggests a post focused on implementing Datadog data visualization for iOS performance monitoring, but no implementation details are provided. Please provide the article body or a readable extraction of the post for a substantive technical summary.

Read original(opens in new tab)