Techlist.io - Korean Tech Blog Curator

figma3 min readCurated summary

Introducing Figma’s New Dev Mode | Figma Blog

Dev Mode is Figma’s dedicated workspace for developers, designed to reduce friction between design and implementation. It provides familiar inspection tools, customizable code output, integrations with development workflows, and an extension for VS Code. Figma’s broader goal is to keep designers and developers working from the same, up-to-date source of truth. ## A Developer-Focused Workspace - Dev Mode works like a browser inspector for Figma files. - Developers can inspect layers and designs to find: - Measurements and implementation specs - Exportable assets - Design-system context - Connections between visual elements and code concepts - The interface is intended to feel familiar to developers who use tools such as Chrome DevTools. ## Faster, More Flexible Code Handoff - Dev Mode’s code panel includes: - A CSS box model - Modern syntax and tree views - Configurable dimension units - Customizable output for different languages and codebases - Generated code is positioned as a starting point that reduces repetitive translation from design to implementation. ## Connecting Design to the Development Toolchain - Figma’s GitHub integration links designs and components with files, issues, and pull requests. - Plugins connect Figma to tools including: - Jira, Linear, and GitHub for project tracking - Storybook for viewing coded components alongside designs - AWS Amplify Studio, Google Relay, and Anima for code generation - Teams can also build custom plugins to match their own workflows. - Design tokens and variables further improve consistency between design systems and code. ## Tracking Work Toward Production - Dev Mode is intended to support the increasingly blurred boundary between design and development. - It helps teams understand what is ready for implementation and maintain context while designs continue to evolve. - Figma emphasizes collaboration around work that remains in progress rather than waiting for designs to be fully polished. ## Bringing Dev Mode into VS Code - The Figma for VS Code extension lets developers: - Review designs - View comments and notifications - Track design changes - Inspect designs without leaving the editor - Use design-based code autocomplete - This keeps design context inside the developer’s primary coding environment. ## Availability and Pricing - Dev Mode and the VS Code extension were announced as beta features, free to all users through the end of 2023. - Beginning in 2024, Dev Mode requires a paid plan. - Dedicated Dev Mode-only access was planned at: - $25 per seat per month on Organization - $35 per seat per month on Enterprise Figma presents Dev Mode as an initial step toward tighter designer-developer collaboration. Teams should use it alongside integrations such as GitHub, Storybook, and VS Code to keep specifications, implementation context, and feedback in one continuously updated workflow.

Read original(opens in new tab)
figma3 min readCurated summary

AI: The Next Chapter in Design | Figma Blog

Figma argues that AI will become a core platform capability reshaping the entire product-development process, not merely another feature. It can accelerate ideation, design, and coding while allowing teams to focus more on problem-solving and creative judgment. Figma announced its acquisition of Diagram as part of this strategy, positioning AI as a force that will change how products are designed, what experiences are created, and who participates in the process. ## Figma’s AI strategy and Diagram acquisition - Figma acquired Diagram, founded by Jordan Singer, whose GPT-3-powered “Designer” plugin generated design concepts from simple prompts. - The acquisition brings Diagram’s team into Figma and builds on Figma’s existing investment in machine learning. - Figma’s open API has already enabled nearly 100 community-built AI plugins. - The company views AI as a platform underlying the entire product-development workflow. ## AI across the product-development process - During discovery, AI could: - Generate and synthesize early ideas from prompts. - Summarize discussions and concepts. - During design, AI could: - Use existing designs and design systems to provide recommendations. - Surface relevant components and patterns. - Help teams produce first drafts faster. - During development, AI could: - Infer design context more effectively. - Generate higher-quality, production-ready code. - The broader goal is to help teams do more work faster while moving their attention toward higher-level problem-solving. ## How design may evolve from pixels to patterns - Design systems already shifted designers away from repetitive details such as border radii and toward composition, direction, and judgment. - Atomic elements such as pixels became reusable components, enabling faster and more consistent workflows. - AI could extend this progression by generating higher-level structures and patterns. - Designers may focus less on assembling basic login components and more on inventing entirely new ways to authenticate. - AI might also recommend color palettes based on a project’s emotional tone or theme. - This could move design beyond familiar interfaces toward smoother, more intuitive, and more human experiences. ## What product teams may design - AI systems such as ChatGPT are shifting interaction away from navigating websites and apps toward asking questions and receiving answers. - AI can reduce the gap between a user’s intention and the actions required to achieve it. - For example, instead of opening a ride-hailing app, entering a destination, comparing options, and requesting a ride, a user could simply say, “Get me to JFK.” - Product builders will need to reconsider whether existing interfaces can deliver the same outcome with fewer steps and decisions. ## The changing role of designers - Technological change has historically transformed design without eliminating the need for thoughtful designers. - Designers have adapted to new platforms, collaborative workflows, and hybrid work. - Figma expects AI to change design roles and collaboration, but frames that shift as an opportunity to spend more time on creative direction, curation, and meaningful problem-solving. The practical recommendation is to treat AI as a foundational design and development capability rather than a standalone feature. Teams should explore how it can remove repetitive work while preserving human judgment, taste, and responsibility for the experiences they create.

Read original(opens in new tab)
figma3 min readCurated summary

Why Roles Are Not Rules | Figma Blog

Figma CTO Kris Rasmussen argues that engineering roles should guide collaboration, not restrict who participates in product decisions. As products become more collaborative and nonlinear, engineers must contribute not only to implementation but also to deciding what to build. However, collaboration works best when teams balance broad input with clear ownership, milestones, and momentum. ## Collaboration Beyond Traditional Roles - Modern product development is increasingly multiplayer, shaped by: - Web technologies such as WebGL and WebAssembly - Collaboration tools like Google Docs - Hybrid and remote work - Engineers are no longer simply responsible for executing plans created by product managers and designers. - At Figma, engineers help determine both: - **How** to build products - **What** to build - Cross-functional collaboration includes working closely with product peers and incorporating customer feedback. ## Early, Open Design Work - Teams are encouraged to gather diverse perspectives at the beginning of a project. - Engineers write down early ideas in concept documents rather than waiting until proposals are polished. - Feedback is collected during the drafting process, allowing designs to evolve collaboratively. - This approach helps expose problems earlier, though it can be difficult to obtain timely feedback from teams focused on their own work. ## Engineering Crits as Feedback Forums - Figma holds regular engineering “crits” across design and engineering organizations. - Crits provide: - Early and frequent feedback - Expert input on technical designs - A dedicated forum for cross-team participation - They are explicitly **not approval meetings**: - No decisions need to be finalized during the session. - Participants identify problems and improve designs without immediately choosing a solution. - Figma uses FigJam so participants can collaborate in real time. - The goal is to improve a proposal until it no longer requires formal approval. ## Balancing Input with Direction - Collaboration can become counterproductive when teams receive too many conflicting ideas. - Excessive feedback may cause projects to: - Lose focus - Get stuck in endless exploration - Struggle with ambiguous tradeoffs, such as defining an initial pricing model - Teams need a balance between: - Diverging to explore possibilities - Converging to make progress - The right balance depends on the organization, culture, product, and desired outcomes. ## Milestones and Momentum - Figma breaks projects into clearly defined milestones. - Milestones help: - Set expectations with stakeholders - Signal when exploration should give way to decisions - Preserve forward momentum - Momentum makes goals feel attainable, while losing momentum can cause teams to question their direction and spin in circles. The practical lesson is to treat roles as areas of responsibility rather than rigid boundaries. Invite collaboration early, use structured forums for feedback, and establish milestones that make it clear when the team must stop exploring and move forward.

Read original(opens in new tab)
figma2 min readCurated summary

Introducing Shortcut letter from the editor | Figma Blog

Shortcut is Figma’s redesigned blog, created to tell richer stories about the people, processes, and ideas behind new products. Its central philosophy mirrors Figma itself: collaboration, experimentation, and the way user input transforms tools over time. The new experience aims to make content more immersive, discoverable, and inspiring rather than static or one-directional. ## A Blog Built Around Collaboration and Play - Figma compares product development to games, where people create “house rules” and shortcuts to improve shared experiences. - This idea reflects Figma’s history as a web-based, multiplayer design tool. - Although collaborative editing initially seemed risky, user behavior and community feedback helped shape Figma’s evolution. - Shortcut will showcase how teams and users bring collaborative design to life through stories about process, craft, and problem-solving. ## A More Immersive Content Experience - The blog redesign replaces a static reading experience with a more dynamic and visually engaging one. - New articles cover: - Design systems - Engineering - Product design and development - Company culture - Behind-the-scenes work - Technical explorations - Examples include stories about Linear responding to a DDoS attack, font failures, and how Figma’s engineers draw inspiration from gaming. ## Improved Navigation and Visual Identity - Readers can sort content by topic and browse curated special collections. - Shortcut features new illustrations and motion work from a range of independent artists. - The redesign treats editorial design, animation, and artwork as essential parts of the storytelling experience. - Figma also invites writers, designers, and other creatives to collaborate with its editorial team. ## Inspiration Through Shared Ideas - Shortcut is built around “input for output”: one idea should lead to another, with inspiration moving in both directions. - The blog aims to encourage readers—and Figma’s own teams—to experiment, share ideas, and have more fun. - Stories range from curated series such as the Future of Design Systems to opinion essays, technical deep dives, and unexpected examples like a band using a Figma slide deck to relaunch its career. Shortcut positions Figma’s blog as an extension of its collaborative product philosophy: a place where stories, disciplines, and perspectives connect to inspire new possibilities.

Read original(opens in new tab)
figma2 min readCurated summary

What’s Happening at Config 2023? | Figma Blog

Config 2023 was covered as a live, informal scene report by Figma editors Herbert Lui and Jenny Xie. The post captures the excitement, humor, attendee reactions, and social-media energy surrounding the event rather than presenting a detailed recap of individual talks. Readers are directed to watch the full talks online and revisit the community conversation under #Config2023 and #confits2023. ## Live Coverage from Config 2023 - Figma’s Editorial Team reported from the event in real time. - The coverage aimed to give people who missed Config a sense of the atmosphere and ongoing activity. - The post included embedded posts from attendees and Figma community members, including jokes about different types of designers attending the conference. ## Community Reactions and Social Media - Attendees shared photos, commentary, and humorous observations on X/Twitter. - The hashtag **#confits2023** highlighted fashion and outfit-related posts from the event. - The social conversation reflected both the professional content and the playful culture surrounding Config. ## Talks and Follow-Up Content - Figma provided a YouTube playlist containing all talks from Config 2023. - Readers were encouraged to catch up on presentations after the event. - The editors closed by celebrating the event’s success, acknowledging “Config withdrawals,” and looking ahead to the following year. The practical takeaway is to use the recorded talk playlist for substantive content, while the blog and accompanying social posts provide a more casual view of Config’s community and atmosphere.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into the platform-level recovery | Datadog (opens in new tab)

Following a massive system-wide outage in March 2023, Datadog successfully restored its EU1 region by identifying that a simple node reboot could resolve network connectivity issues caused by a faulty system patch. While the team managed to restore 100 percent of compute capacity within hours, the recovery effort was subsequently hindered by cloud provider infrastructure limits and IP address exhaustion. This post-mortem highlights the complexities of scaling hierarchical Kubernetes environments under extreme pressure and the importance of accounting for "black swan" capacity requirements. ## Hierarchical Kubernetes Recovery Datadog utilizes a strict hierarchy of Kubernetes clusters to manage its infrastructure, which necessitated a granular, three-tiered recovery approach. Because the outage affected network connectivity via `systemd-networkd`, the team had to restore components in a specific order to regain control of the environment. * **Parent Control Planes:** Engineers first rebooted the virtual machines hosting the parent clusters, which manage the control planes for all other clusters. * **Child Control Planes:** Once parent clusters were stable, the team restored the control planes for application clusters, which run as pods within the parent infrastructure. * **Application Worker Nodes:** Thousands of worker nodes across dozens of clusters were restarted progressively to avoid overwhelming the control planes, reaching full capacity by 12:05 UTC. ## Scaling Bottlenecks and Cloud Quotas Once the infrastructure was online, the team attempted to scale out rapidly to process a massive backlog of buffered data. This surge in demand triggered previously unencountered limitations within the Google Cloud environment. * **VPC Peering Limits:** At 14:18 UTC, the platform hit a documented but overlooked limit of 15,500 VM instances within a single network peering group, blocking all further scaling. * **Provider Intervention:** Datadog worked directly with Google Cloud support to manually raise the peering group limit, which allowed scaling to resume after a nearly four-hour delay. ## IP Address and Subnet Capacity Even after cloud-level instance quotas were lifted, specific high-traffic clusters processing logs and traces hit a secondary bottleneck related to internal networking. * **Subnet Exhaustion:** These clusters attempted to scale to more than twice their normal size, quickly exhausting all available IP addresses in their assigned subnets. * **Capacity Planning Gaps:** While Datadog typically targets a 66% maximum IP usage to allow for a 50% scale-out, the extreme demands of the recovery backlog exceeded these safety margins. * **Impact on Backlog:** For six hours, the lack of available IPs forced these clusters to process data significantly slower than the rest of the recovered infrastructure. ## Recovery Summary The EU1 recovery demonstrates that even when hardware is functional, software-defined limits can create cascading delays. Organizations should not only monitor their own resource usage but also maintain visibility into cloud provider quotas and ensure that subnet allocations account for extreme recovery scenarios where workloads may need to double or triple in size momentarily.

datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level recovery

Datadog’s March 8, 2023 outage removed 60% of its compute capacity, forcing teams to restore infrastructure in stages while accounting for regional and cloud-provider differences. In EU1, recovery depended on rebooting affected nodes, restoring Kubernetes control planes in a strict hierarchy, and gradually bringing application capacity back online. Scaling afterward exposed infrastructure limits that had not been considered during normal operations. ## EU1 Platform Recovery - A system patch disconnected affected EU1 nodes from the network, but the nodes could be recovered through reboots. - Recovery was initially slowed by the lack of observability and unavailable Kubernetes APIs. - Datadog operates: - **Parent clusters**, which host the control-plane pods for other clusters. - **Child clusters**, where Datadog applications run. - This hierarchy allows Datadog to use Kubernetes deployment, replacement, rolling-update, and autoscaling capabilities for child-cluster control planes. - Parent-cluster control planes run on VMs and are managed with `systemd`. ## Restoring Kubernetes Clusters Because both parent and child environments were affected by the Ubuntu 22.04 issue, recovery had to follow a strict sequence: - **Parent control planes:** Nodes running Cilium were rebooted to restore network connectivity. This finished by 08:45 UTC. - **Child control planes:** All parent-cluster nodes hosting child control-plane pods were rebooted. This finished by 09:30 UTC. - **Application nodes:** Thousands of instances across dozens of child clusters were restarted. - Recovery reached 60% by 10:20 UTC. - All application nodes were restored by 12:05 UTC. - Restarts were prioritized by workload importance and paced to avoid overwhelming Kubernetes control planes. ## Scaling Capacity and Recovering Backlogs After restoring the clusters, Datadog needed substantial additional capacity to process data buffered during the outage. - EU1 hit a Google Cloud mesh limit of **15,500 VM instances** at 14:18 UTC. - Instance creation failures became apparent around 15:00 UTC. - Datadog had not checked this documented limit before the incident, but Google Cloud quickly raised it after Datadog submitted a high-priority request. - Autoscaling also exhausted the IP capacity of subnets used by three log- and trace-processing clusters. - These clusters normally used about 35–45% of their IP capacity, but the backlog caused autoscaling to request more than twice their usual replica counts, filling the subnets. ## Practical Lessons The recovery demonstrated that restoring compute capacity is not enough: teams must also understand dependency order, control-plane architecture, cloud-provider quotas, and network-address limits. Capacity planning should account for severe backlog-driven scaling, not just normal operating utilization, and documented infrastructure limits should be validated before emergencies occur.

Read original(opens in new tab)
datadog3 min readCurated summary

2023-03-08 incident: A deep dive into our incident response

Datadog’s March 8, 2023 global outage tested an incident-response process designed for large-scale failures. The company’s monitoring, on-call structure, training, and blameless culture enabled a coordinated response, but the incident also exposed challenges in diagnosing and managing a rapidly evolving, global outage. Datadog’s central lesson is that effective response depends less on rigid runbooks than on preparation, clear ownership, autonomous decision-making, and continuous learning. ## Datadog’s Incident Response Model - Datadog follows a “you build it, you own it” operating model. - Teams instrument their services extensively and configure monitors to detect problems around the clock. - Independent, out-of-band monitoring checks Datadog’s APIs from outside its infrastructure, ensuring that monitoring still works if Datadog itself becomes unavailable. - Slack channels are automatically created for incidents to provide shared situational awareness and enable additional engineers to contribute. ## Handling High-Severity Incidents - Senior engineers rotate on call for incidents involving substantial customer impact or multiple teams. - The first senior responder becomes the incident commander and retains overall responsibility. - A communications lead may manage internal updates and coordination. - For the most serious incidents, an engineering executive and customer-support manager join to provide leadership, business context, and customer-facing communication. - The incident commander remains accountable for coordinating the overall response. ## Preparation, Training, and Postmortems - Datadog uses a relatively low threshold for declaring incidents, giving engineers frequent practice with its response process. - Engineers complete incident-response training before joining an on-call rotation and repeat refresher training every six months. - Training covers on-call responsibilities, response roles, and blameless investigation practices. - Every high-severity incident receives a detailed postmortem focused on preventing recurrence. - Automation prompts responders to begin postmortems while the incident is still fresh. ## Autonomy and a Blameless Culture - Because large systems change constantly, detailed recovery procedures can quickly become outdated. - Datadog therefore gives engineers authority to choose the best response based on their knowledge of the affected services. - The company treats failures as weaknesses in systems rather than evidence of individual fault. - Blamelessness is intended to encourage creativity, honesty, and effective decision-making under pressure. ## The March 8 Outage - A systemd upgrade began around 06:00 UTC and ultimately triggered the outage. - Monitoring detected the problem within three minutes, and engineering teams were paged shortly afterward. - A high-severity incident was declared at 06:18, with an incident commander joining five minutes later. - The first public status update was posted at 06:31, and the outage was officially diagnosed as global at 06:32. - By 07:20, responders identified a Kubernetes failure and unhealthy intake systems as central problems. - Engineers confirmed by 08:00 that the Kubernetes failure was not spreading to additional or newly provisioned nodes. - A working mitigation for the EU1 region was found by 08:30. - Most US1 compute capacity recovered automatically by 11:00, while teams began organizing a longer recovery effort. - At 11:36, unattended upgrades were identified as the triggering event. - Compute capacity in EU1—the first step toward recovery—was restored by 12:05. ## Practical Lessons Datadog’s experience demonstrates the value of independent monitoring, practiced incident roles, rapid communication, and empowered responders. Organizations operating complex systems should regularly rehearse incident management, invest in resilient observability outside the primary platform, and use blameless postmortems to turn major outages into improvements.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into our incident response | Datadog (opens in new tab)

Datadog’s first global outage on March 8, 2023, served as a rigorous stress test for their established incident response framework and "you build it, you own it" philosophy. While the outage was triggered by a systemic failure during a routine systemd upgrade, the company's commitment to blameless culture and decentralized engineering autonomy allowed hundreds of responders to coordinate a complex recovery across multiple regions. Ultimately, the event validated their investment in out-of-band monitoring and rigorous, bi-annual incident training as essential components for managing high-scale system disasters. ## Incident Response Structure and Philosophy * Datadog employs a decentralized "you build it, you own it" model where individual engineering teams are responsible for the 24/7 health and monitoring of the services they build. * For high-severity incidents, a specialized rotation is paged, consisting of an Incident Commander to lead the response, a communications lead, and a customer liaison to manage external messaging. * The organization prioritizes "people over process," empowering engineers to use their judgment to find creative solutions rather than following rigid, pre-written playbooks that may not apply to unprecedented failures. * A blameless culture is strictly maintained across all levels of the company, ensuring that post-incident investigations focus on systemic improvements rather than assigning fault to individuals. ## Multi-Layered Monitoring Strategy * Standard telemetry provides internal visibility, but Datadog also maintains "out-of-band" monitoring that operates completely outside its own infrastructure. * This out-of-band system interacts with Datadog APIs exactly like a customer would, ensuring that engineers are alerted even if the internal monitoring platform itself becomes unavailable. * Communication is streamlined through a dedicated Slack incident app that automatically generates coordination channels, providing situational awareness to any engineer who joins the effort. ## Anatomy of the March 8 Outage * The outage began at 06:00 UTC, triggered by a systemd upgrade that caused widespread Kubernetes failures and prevented pods from restarting correctly. * The global nature of the outage was diagnosed within 32 minutes of the initial monitoring alerts, leading to the activation of executive on-calls and the customer support management team. * Responders identified "unattended upgrades" as the incident trigger approximately five and a half hours after the initial failure. * Recovery was executed in stages: compute capacity was restored first in the EU1 region, followed by the US1 region, with full infrastructure restoration completed by 19:00 UTC. Organizations should treat incident response as a perishable skill that requires constant practice through a low threshold for declaring incidents and regular training. By combining out-of-band monitoring with a culture that empowers individual engineers to act autonomously during a crisis, teams can more effectively navigate the "not if, but when" reality of large-scale system failures.

datadog2 min readCurated summary

Not just another network latency issue: How we unraveled a series of hidden bottlenecks | Datadog

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content primarily consists of Datadog’s product navigation rather than the full blog post, so it does not include Gartner’s evaluation criteria, Datadog’s specific strengths, or comparative analysis. ## Gartner Recognition - Datadog is identified as a **Leader** in Gartner’s **Magic Quadrant for Observability Platforms 2026**. - The announcement links to a Gartner-related resource hosted by Datadog. - No details are provided in the supplied text about: - Gartner’s assessment of Datadog’s execution or vision - The other vendors evaluated - Specific capabilities that led to the ranking - Any limitations or cautions noted by Gartner ## Datadog’s Observability Offering The surrounding navigation shows that Datadog positions observability as part of a broad platform covering: - Infrastructure and cloud monitoring - Application performance monitoring and profiling - Logs, metrics, databases, and data pipelines - Real user, synthetic, and mobile monitoring - Security and cloud protection - CI/CD and software delivery - Incident response, service management, and automation - AI-powered investigation and agent observability The supplied content does not explain how these products are integrated or how they support the Leader designation. Overall, the material communicates Datadog’s recognition as a Gartner observability-platform Leader, but the full article or report is needed for a meaningful technical summary.

Read original(opens in new tab)
datadog3 min readCurated summary

Not just another network latency issue: How we unraveled a series of hidden bottlenecks

Repeated high-startup-latency pages in Datadog’s usage estimation service were caused by several independent bottlenecks rather than application changes. The investigation eventually identified four issues: CPU-throttled Envoy sidecars, a Linux kernel bug affecting ENA transmit queues, insufficient EC2 network bandwidth, and requests routed to terminating cache pods. Fixing each layer progressively reduced remote-cache p99 latency from roughly one second to its normal level of about 100 ms. ## Service Architecture and the Original Symptoms - The service consists of router, counter, and aggregator applications. - At startup, `counter` loads data from a remote cache into a local cache. - While the local cache is populating, request processing is slower and backlog grows. - Normal p99 remote-cache latency was approximately 100 ms, but it exceeded one second during deployments. - Scaling the remote cache did not help, indicating that the cache itself was not underprovisioned. ## CPU-Throttled Envoy Sidecars - Requests to the remote cache passed through an Envoy sidecar that batched queries into packets. - When `counter` restarted, Envoy reached its two-core CPU limit and was throttled. - Delayed request and response processing caused retries, TCP retransmits, and increased remote-cache latency. - Increasing Envoy’s CPU allocation eliminated the issue in staging and reduced production latency, but did not fully resolve rollout spikes. ## Linux Kernel and ENA Transmit-Queue Bug - Investigation of system and network metrics revealed a Linux kernel bug affecting AWS Elastic Network Adapter traffic. - The kernel mapped all outbound traffic to the first transmit queue instead of distributing it across eight queues. - This limited throughput and caused retransmits during high-traffic periods such as deployments. - A hotfix distributed traffic across all eight queues. - The change removed non-rollout latency spikes but rollout latency still fluctuated between 200 and 600 ms. ## EC2 Network Bandwidth Limits - ENA metrics showed that instances exceeded AWS inbound and outbound bandwidth allowances. - AWS dropped packets at the hypervisor when those limits were exceeded, causing retransmissions and slower cache requests. - Migrating to network-optimized EC2 instance types with higher bandwidth allowances largely restored p99 latency to around 100 ms. - Occasional one-second spikes continued despite the improvement. ## Requests Sent to Terminating Cache Pods - Remaining spikes correlated with remote-cache pods that were shutting down. - Clients continued sending requests to terminating pods, leading to one-second timeouts and retries. - The cache’s graceful-shutdown behavior did not adequately wait for Envoy clients’ in-flight requests. - The team added a `preStop` hook that sets an `XXX_MAINTENANCE_MODE` key to notify clients before termination and began coordinating shutdown with outstanding requests. The incident demonstrates the importance of tracing latency across the entire request path, from application startup through proxies, kernel networking, hardware interfaces, cloud bandwidth limits, and pod lifecycle behavior. Layered system metrics and component-level investigation were necessary to eliminate alert fatigue and restore reliable deployment behavior.

Read original(opens in new tab)
figma4 min readCurated summary

How a Figma Slide Deck Helped Indie Rock Duo Tanlines Launch a Comeback | Figma | Figma Blog

Tanlines’ comeback album, *The Big Mess*, grew out of Jesse Cohen and Eric Emm’s changing lives as parents, professionals, and musicians. To promote the album’s first single, “Outer Banks,” they turned a corporate presentation deck into a music video, using Figma and Google Meet to parody workplace communication while expressing the song’s emotional themes. The project shows how familiar professional tools can become creative, collaborative artistic formats. ## Returning to Tanlines - Cohen and Emm have made music together since 2008, evolving from experimental electronic work and remixes into pop music. - After the band’s second album, Cohen stepped away to raise his first child and later made a children’s record. - Emm eventually assembled songs that became the foundation for *The Big Mess*, and Cohen helped finish them during a working trip to rural Connecticut. - Cohen now views Tanlines as an ongoing life project, similar to the documentary series *7 Up*, revisiting the band as their circumstances change. ## Turning a Corporate Deck into a Music Video - Cohen’s years at YouTube Music and Nike introduced him to decks and meetings as a dominant form of professional communication. - He wanted to tell the story of “Outer Banks” using that shared workplace language. - Former coworker Hunter Ellenbarger created the presentation in Figma. - The original idea was to make a deck for every song and release them through Google Drive or a Figma folder. - Hunter’s deck was strong enough to become the centerpiece of the video, with Emm singing the lyrics as though he were reading the presentation aloud. - Presenting the Figma deck through Google Meet reinforced the joke and connected the video to remote-work culture. ## Making the Deck Feel Authentically Corporate - Cohen’s initial design leaned toward Nike’s polished, editorial presentation style. - Ellenbarger deliberately made the visuals more generic, using: - Graphs and arrows - Flow charts - Corporate phrases such as “insights,” “year over year,” and “results” - Inspirational quotes and familiar meeting-slide conventions - The overly recognizable format creates both humor and emotional resonance for viewers accustomed to workplace presentations. ## The Emotional Meaning of “Outer Banks” - The song references North Carolina’s Outer Banks, a place Cohen has visited throughout his life. - Rather than presenting the beach as a carefree party setting, the song explores memory, family, and mixed emotions. - Its upbeat sound carries a melancholy undertone, consistent with Tanlines’ recurring balance of light and darkness. - The video reframes an ordinary corporate tool as a medium for music, art, and humor, allowing professional audiences to feel recognized rather than mocked. ## Building a New Visual Identity - Longtime collaborator Teddy Blanks helped shape the album packaging, promotional poster, and merchandise. - The visual material draws on 1950s photographs taken in Greece by Emm’s wife’s grandfather while documenting folk tales and songs. - The new artwork moves away from the band’s earlier practice of featuring the musicians’ faces on album covers. - Blanks’ simple, clear design choices support the sense that *The Big Mess* represents a new era for the band. ## Embracing Change and Uncertainty - Cohen describes “aging gracefully” as adapting creative expression to new responsibilities and circumstances. - Returning to music was unexpected but allowed him to reconnect with an important part of his identity. - Unlike earlier album cycles, the band’s future is less predictable, with more uncertainty around touring, work, family, and sustaining a music career in middle age. The project suggests that creative work does not need to reject familiar professional formats; it can transform them. By combining Figma’s collaborative presentation tools, Google Meet’s remote setting, and the language of corporate decks, Tanlines made a comeback video that was both a workplace satire and a sincere reflection on adult life.

Read original(opens in new tab)
datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level impact

Datadog’s March 8, 2023 outage was caused by an unexpected interaction between Ubuntu 22.04, systemd-networkd, and an automated security patch. A systemd change introduced behavior that flushed unfamiliar IP routing rules whenever systemd-networkd restarted; a CVE patch triggered that restart across many hosts. Because the patch was installed automatically and outside Datadog’s carefully staged deployment process, infrastructure across regions and cloud providers was affected simultaneously. ## A Systemd Behavior Change - systemd v248 introduced a systemd-networkd startup behavior that removed IP rules it did not recognize. - systemd v249 added the `ManageForeignRoutingPolicyRules` setting, which could disable this behavior, but the default configuration continued managing foreign rules. - These changes were backported to older systemd releases. - Ubuntu 20.04 used systemd v245, which did not flush IP rules during a systemd-networkd restart. - Ubuntu 22.04, adopted progressively by Datadog beginning in November 2022, used systemd v249 with the behavior enabled. Initially, the change caused no visible problems because systemd-networkd generally started only when new hosts were created, before Datadog’s custom routing rules existed. ## The Security Patch That Triggered the Problem - On March 7, 2023, Ubuntu released systemd patch `249.11-0ubuntu3.7` for a CVE. - Installing the patch restarted all systemd components, including systemd-networkd. - That restart caused systemd-networkd to flush routing policy rules on affected Ubuntu 22.04 hosts. - Ubuntu 20.04 hosts received a similar patch but were not affected because systemd v245 did not exhibit the problematic restart behavior. ## Unattended Upgrades Created Broad Exposure Datadog used Ubuntu’s default unattended-upgrade configuration: - Package metadata was downloaded twice daily using `apt-daily.timer`, with randomized delays of up to 12 hours. - Upgrades ran daily using `apt-daily-upgrade.timer`, between 06:00 and 07:00 UTC. - Only security updates and required dependencies were automatically installed. - Regular updates from the `-updates` repository were excluded. This meant many hosts automatically installed the systemd security patch during the same daily upgrade window. Not every host was affected: more than 90% of the fleet used Ubuntu 22.04, and some nodes had not yet downloaded the patch when their upgrade ran. ## Conflict with Datadog’s Deployment Process - Datadog normally updates nodes by replacing them automatically rather than relying on unattended upgrades. - Its standard process validates changes on experimental clusters, then progressively deploys them through staging and production. - Deployments are normally limited to selected clusters, availability zones, and regions before expanding. - The unattended systemd patch bypassed this process because it was installed directly on existing hosts. - As a result, a low-level networking change propagated across otherwise isolated regions and cloud providers at nearly the same time. Datadog’s experience demonstrates that even security-only automated updates can introduce coordinated infrastructure risk. Critical system packages should be tested and rolled out through the same staged process as other production changes, or their automated upgrades should be carefully constrained and monitored.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into the platform-level impact | Datadog (opens in new tab)

The March 2023 Datadog outage was triggered by a simultaneous, global failure across multiple cloud providers and regions, caused by an unexpected interaction between a systemd security patch and Ubuntu 22.04’s default networking behavior. While Datadog typically employs rigorous, staged rollouts for infrastructure changes, the automated nature of OS-level security updates bypassed these controls. The incident highlights the hidden risks in system-level defaults and the potential for "unattended upgrades" to create synchronized failures across supposedly isolated environments. ## The systemd-networkd Routing Change * In December 2020, systemd version 248 introduced a change where `systemd-networkd` flushes all IP routing rules it does not recognize upon startup. * Version 249 introduced the `ManageForeignRoutingPolicyRules` setting, which defaults to "yes," confirming this management behavior for any rules not explicitly defined in systemd configuration files. * These changes were backported to earlier versions (v247 and v248) but were notably absent from v245, the version used in Ubuntu 20.04. ## Dormant Risks in the Ubuntu 22.04 Migration * Datadog began migrating its fleet from Ubuntu 20.04 to 22.04 in late 2022, eventually reaching 90% coverage across its infrastructure. * Ubuntu 22.04 utilizes systemd v249, meaning the majority of the fleet was susceptible to the routing rule flushing behavior. * The risk remained dormant during the initial rollout because `systemd-networkd` typically only starts during the initial boot sequence when no complex routing rules have been established yet. ## The Trigger: Unattended Upgrades and the CVE Patch * On March 7, 2023, a security patch for a systemd CVE was released to the Ubuntu security repositories. * Datadog’s fleet used the Ubuntu default configuration for `unattended-upgrades`, which automatically installs security-labeled patches once a day, typically between 06:00 and 07:00 UTC. * The installation of the patch forced a restart of the `systemd-networkd` service on active, running nodes. * Upon restarting, the service identified existing IP routing rules (crucial for container networking) as "foreign" and deleted them, effectively severing network connectivity for the nodes. ## Failure of Regional Isolation * Because the security patch was released globally and the automated upgrade window was synchronized across regions, the failure occurred nearly simultaneously worldwide. * This automation bypassed Datadog’s standard practice of "baking" changes in staging and experimental clusters for weeks before proceeding to production. * Nodes on the older Ubuntu 20.04 (systemd v245) were unaffected by the patch, as that version of systemd does not flush IP rules upon a service restart. To mitigate similar risks, infrastructure teams should consider explicitly disabling the management of foreign routing rules in systemd-networkd configuration when using third-party networking plugins. Furthermore, while automated security patching is a best practice, organizations must balance the speed of patching with the need for controlled, staged rollouts to prevent global configuration drift or synchronized failures.

figma2 min readCurated summary

Four years in, here’s what Config tells us about the state of design | Figma Blog

Config 2023 submissions suggest that the design industry is emerging from pandemic-era uncertainty with renewed optimism. Figma analyzed more than 1,000 conference proposals and found increasing interest in accessibility, creativity enabled by design systems, and broader collaboration. Overall, the submissions portray designers as focused on turning recent challenges into opportunities for more scalable, inclusive, and effective product development. ## A More Optimistic Design Community - Config submissions increased from 420 in 2021 to 520 in 2022 and more than 1,000 in 2023. - Positive sentiment rose from 63% of submissions in 2021 to 72% in 2023. - Submissions increasingly emphasized possibilities and ways teams had “thrived,” rather than focusing on outdated practices or persistent problems. - References to the pandemic fell to one-sixth of their 2021 level. - Mentions of “remote” declined by roughly 20%, suggesting that teams are adapting to changed work patterns. ## Design Systems as Creative Infrastructure - The perceived divide between creative freedom and systematic processes is narrowing. - Designers increasingly view design systems as tools that reduce repetitive work and create more space for creative thinking. - Teams operating at scale are investing in design tokens and reusable systems to maintain speed and consistency. - In 2021, design-system discussions focused mainly on foundational tasks such as auditing, scaling, and establishing basic infrastructure. - By 2023, submissions connected design systems more directly with creativity and visual expression, using terms such as “art,” “transition,” “color,” and “creating.” The emerging view is that systems do not suppress creativity; they can provide the structure and efficiency needed to support it.

Read original(opens in new tab)