Kubernetes

144 posts

datadog3 min readCurated summary

Engineering VP spotlight: Ivo Dimitrov

Ivo Dimitrov’s career evolved from low-level systems programming into engineering leadership focused on large-scale distributed storage. His experience at Microsoft and LinkedIn shaped his approach to building scalable data platforms, while Datadog attracted him with its talented people, modern technology, and culture of experimentation. Today, he leads Datadog’s Distributed Data Systems organization, supporting the company’s metrics, events, query, alerting, and analytics infrastructure. ## From Electrical Engineering to Systems Programming - Dimitrov initially studied electrical engineering and became interested in software while working on digital control systems. - His early work included contributing to a real-time operating system kernel. - He spent roughly a decade as an individual contributor, specializing in: - High-performance systems - Low-level programming - C and C++ - System software ## Transition from Individual Contributor to Manager - At Microsoft, Dimitrov worked on an early version of Azure Blob Storage. - Following a reorganization, he accepted an opportunity to lead his team despite having no prior management experience. - Microsoft supported the transition through: - Leadership mentorship - Formal management training - Guidance on communication, conflict resolution, and interpersonal leadership - He discovered that management allowed him to expand his ownership beyond individual projects and influence broader organizational outcomes. - The role combined his technical background with responsibilities such as cross-functional coordination, team development, and engineering strategy. ## Building Internet-Scale Storage at Microsoft and LinkedIn - At Microsoft, Dimitrov worked on storage systems supporting Hotmail. - After joining LinkedIn in 2014, he adapted to a technology environment centered on open source tools such as MySQL and Java. - He led development of Espresso, LinkedIn’s proprietary key-value storage platform. - The platform matured into a core system supporting approximately 95 percent of LinkedIn’s data sets. - He also helped oversee several other large-scale storage projects: - **Venice**, an open source platform for serving derived data - **Ambry**, an open source blob storage system - **Helix**, an open source cluster manager - These systems supported critical parts of LinkedIn’s internet-scale infrastructure. ## Why Datadog Was Appealing - Dimitrov was drawn to Datadog by three main factors: - Highly capable engineers and leaders - Interesting, modern technology - The opportunity to contribute to a rapidly growing company - Compared with the legacy systems and processes that had accumulated at LinkedIn, Datadog offered less bureaucracy and more freedom to: - Take thoughtful risks - Experiment - Deliver quickly - Fail fast and learn - Iterate and innovate - He was particularly interested in Datadog’s Kubernetes-based Metrics and Events platforms and the challenge of building best-in-class infrastructure during the company’s growth. ## Distributed Data Systems at Datadog - Dimitrov leads the Distributed Data Systems organization, which owns a portfolio of storage and data technologies. - Its responsibilities include: - **Metrics**, supporting metrics and time-series data - **Events**, handling semi-structured data such as logs, profiles, and traces - **Driveline**, a main-memory database optimized for online analytics - The **Cross-Platform Queries** team provides a unified query interface across systems that historically exposed separate, domain-specific APIs. - This reduces the learning curve for engineers and customers. - It abstracts the underlying data stores behind a common API. - The organization also operates Datadog’s Alerts platform, which generates a large share of the queries sent to the Metrics and Events systems. Dimitrov’s experience demonstrates how deep systems expertise can translate into effective engineering leadership. His recommendation by example is to remain technically engaged while expanding one’s scope—from writing individual components to shaping teams, platforms, and long-term engineering direction.

Read original(opens in new tab)
datadog1 min readCurated summary

.NET Continuous Profiler: Memory usage | Datadog

Datadog is presented as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. However, the provided content contains only the announcement headline, link, and website navigation; it does not include the blog post’s analysis, criteria, or supporting details. ## Announcement - Datadog’s headline claim is recognition as a “Leader” in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked page appears to be a Datadog resource or announcement page. ## Available Product Scope The navigation indicates that Datadog’s observability platform spans: - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless systems, and cloud costs - Application performance monitoring, profiling, and dynamic instrumentation - Logs, databases, data pipelines, and data quality - Security monitoring and cloud security - Real user monitoring, synthetic monitoring, session replay, and error tracking - CI/CD visibility, testing, developer portals, and software delivery - Incident management, service catalogs, SLOs, workflow automation, and AI-powered investigation Because the actual article text is missing, no further claims about Gartner’s evaluation or Datadog’s strengths can be reliably summarized.

Read original(opens in new tab)
datadog1 min readCurated summary

How we brought Datadog's data visualization to iOS: A focus on performance | Datadog

The provided content does not include the blog post’s actual article text. It contains Datadog’s navigation menu, product links, and a reference to an engineering post about bringing Datadog data visualization to iOS performance, so its technical argument and conclusion cannot be reliably summarized. ## Available Content - Datadog promotes its recognition as a Leader in the Gartner Magic Quadrant for Observability Platforms. - The navigation lists products across: - Infrastructure and application monitoring - Logs, databases, and data observability - Security - Digital experience monitoring - Software delivery - Service management - AI and platform capabilities - The referenced engineering URL suggests a post focused on implementing Datadog data visualization for iOS performance monitoring, but no implementation details are provided. Please provide the article body or a readable extraction of the post for a substantive technical summary.

Read original(opens in new tab)
datadog1 min readCurated summary

Engineering spotlight: Marie-Laure Bardonnet | Datadog

The provided content does not include the actual blog post. It contains Datadog’s navigation menu and a promotional banner announcing its recognition as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no article text to summarize. ## Datadog’s Observability Announcement - Datadog promotes its recognition as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked announcement appears to be a marketing resource rather than the blog post itself. ## Datadog Product Areas Listed - Infrastructure monitoring, metrics, containers, Kubernetes, networking, serverless, and cloud costs - Application performance monitoring, profiling, dynamic instrumentation, and database monitoring - Log management, observability pipelines, and sensitive data scanning - Security, including cloud security, SIEM, workload protection, and code security - Digital experience tools such as RUM, session replay, synthetic monitoring, and error tracking - Software delivery, CI visibility, testing, feature flags, and code coverage - Service management, incident response, workflows, dashboards, and AI capabilities The article body or source text is needed for a meaningful section-by-section summary.

Read original(opens in new tab)
datadog1 min readCurated summary

Scaling Self-Serve Analytics: The Tools Empowering 5,000 Employees | Datadog

The provided content does not include the blog post itself. It contains Datadog’s navigation menu and a promotional banner announcing its position as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no technical discussion from the article. ## Available Content - The page appears to be a Datadog engineering blog post related to a CrunchConf talk on self-serve analytics. - The supplied text primarily lists Datadog products across: - Infrastructure and application monitoring - Data and log management - Security - Digital experience - Software delivery - Service management - AI capabilities - A separate banner promotes Datadog’s recognition by Gartner. ## Missing Article Details - No article title, introduction, body sections, technical examples, architecture, or conclusions are included. - The available text is insufficient to accurately summarize the post’s arguments or implementation details. Please provide the blog post’s main text or a complete extraction of the page for a substantive summary.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering spotlight: Jeromy Carriere | Datadog

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content does not include the report’s evaluation criteria or detailed rationale, but it presents Datadog as a broad observability platform spanning infrastructure, applications, data, logs, security, digital experience, software delivery, and AI. ## Gartner Recognition - Datadog highlights its position as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked resource appears to provide the full Gartner-related announcement and assessment. ## Broad Observability Coverage - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database monitoring, data-stream and job monitoring, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, vulnerability, compliance, SIEM, and application/API protection. - **Digital experience:** real-user monitoring, session replay, synthetic monitoring, product analytics, mobile testing, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, incident response, SLOs, workflow automation, and developer portals. - **AI capabilities:** AI agents, investigation tools, GPU monitoring, integrations, MCP support, and AI-assisted chat and coding. Datadog’s positioning rests on consolidating telemetry and operational workflows across the technology stack. To understand the Gartner recognition in depth, readers would need the linked report, since the supplied excerpt contains the announcement and product navigation but not the supporting analysis.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into the platform-level recovery | Datadog (opens in new tab)

Following a massive system-wide outage in March 2023, Datadog successfully restored its EU1 region by identifying that a simple node reboot could resolve network connectivity issues caused by a faulty system patch. While the team managed to restore 100 percent of compute capacity within hours, the recovery effort was subsequently hindered by cloud provider infrastructure limits and IP address exhaustion. This post-mortem highlights the complexities of scaling hierarchical Kubernetes environments under extreme pressure and the importance of accounting for "black swan" capacity requirements. ## Hierarchical Kubernetes Recovery Datadog utilizes a strict hierarchy of Kubernetes clusters to manage its infrastructure, which necessitated a granular, three-tiered recovery approach. Because the outage affected network connectivity via `systemd-networkd`, the team had to restore components in a specific order to regain control of the environment. * **Parent Control Planes:** Engineers first rebooted the virtual machines hosting the parent clusters, which manage the control planes for all other clusters. * **Child Control Planes:** Once parent clusters were stable, the team restored the control planes for application clusters, which run as pods within the parent infrastructure. * **Application Worker Nodes:** Thousands of worker nodes across dozens of clusters were restarted progressively to avoid overwhelming the control planes, reaching full capacity by 12:05 UTC. ## Scaling Bottlenecks and Cloud Quotas Once the infrastructure was online, the team attempted to scale out rapidly to process a massive backlog of buffered data. This surge in demand triggered previously unencountered limitations within the Google Cloud environment. * **VPC Peering Limits:** At 14:18 UTC, the platform hit a documented but overlooked limit of 15,500 VM instances within a single network peering group, blocking all further scaling. * **Provider Intervention:** Datadog worked directly with Google Cloud support to manually raise the peering group limit, which allowed scaling to resume after a nearly four-hour delay. ## IP Address and Subnet Capacity Even after cloud-level instance quotas were lifted, specific high-traffic clusters processing logs and traces hit a secondary bottleneck related to internal networking. * **Subnet Exhaustion:** These clusters attempted to scale to more than twice their normal size, quickly exhausting all available IP addresses in their assigned subnets. * **Capacity Planning Gaps:** While Datadog typically targets a 66% maximum IP usage to allow for a 50% scale-out, the extreme demands of the recovery backlog exceeded these safety margins. * **Impact on Backlog:** For six hours, the lack of available IPs forced these clusters to process data significantly slower than the rest of the recovered infrastructure. ## Recovery Summary The EU1 recovery demonstrates that even when hardware is functional, software-defined limits can create cascading delays. Organizations should not only monitor their own resource usage but also maintain visibility into cloud provider quotas and ensure that subnet allocations account for extreme recovery scenarios where workloads may need to double or triple in size momentarily.

datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level recovery

Datadog’s March 8, 2023 outage removed 60% of its compute capacity, forcing teams to restore infrastructure in stages while accounting for regional and cloud-provider differences. In EU1, recovery depended on rebooting affected nodes, restoring Kubernetes control planes in a strict hierarchy, and gradually bringing application capacity back online. Scaling afterward exposed infrastructure limits that had not been considered during normal operations. ## EU1 Platform Recovery - A system patch disconnected affected EU1 nodes from the network, but the nodes could be recovered through reboots. - Recovery was initially slowed by the lack of observability and unavailable Kubernetes APIs. - Datadog operates: - **Parent clusters**, which host the control-plane pods for other clusters. - **Child clusters**, where Datadog applications run. - This hierarchy allows Datadog to use Kubernetes deployment, replacement, rolling-update, and autoscaling capabilities for child-cluster control planes. - Parent-cluster control planes run on VMs and are managed with `systemd`. ## Restoring Kubernetes Clusters Because both parent and child environments were affected by the Ubuntu 22.04 issue, recovery had to follow a strict sequence: - **Parent control planes:** Nodes running Cilium were rebooted to restore network connectivity. This finished by 08:45 UTC. - **Child control planes:** All parent-cluster nodes hosting child control-plane pods were rebooted. This finished by 09:30 UTC. - **Application nodes:** Thousands of instances across dozens of child clusters were restarted. - Recovery reached 60% by 10:20 UTC. - All application nodes were restored by 12:05 UTC. - Restarts were prioritized by workload importance and paced to avoid overwhelming Kubernetes control planes. ## Scaling Capacity and Recovering Backlogs After restoring the clusters, Datadog needed substantial additional capacity to process data buffered during the outage. - EU1 hit a Google Cloud mesh limit of **15,500 VM instances** at 14:18 UTC. - Instance creation failures became apparent around 15:00 UTC. - Datadog had not checked this documented limit before the incident, but Google Cloud quickly raised it after Datadog submitted a high-priority request. - Autoscaling also exhausted the IP capacity of subnets used by three log- and trace-processing clusters. - These clusters normally used about 35–45% of their IP capacity, but the backlog caused autoscaling to request more than twice their usual replica counts, filling the subnets. ## Practical Lessons The recovery demonstrated that restoring compute capacity is not enough: teams must also understand dependency order, control-plane architecture, cloud-provider quotas, and network-address limits. Capacity planning should account for severe backlog-driven scaling, not just normal operating utilization, and documented infrastructure limits should be validated before emergencies occur.

Read original(opens in new tab)
datadog3 min readCurated summary

2023-03-08 incident: A deep dive into our incident response

Datadog’s March 8, 2023 global outage tested an incident-response process designed for large-scale failures. The company’s monitoring, on-call structure, training, and blameless culture enabled a coordinated response, but the incident also exposed challenges in diagnosing and managing a rapidly evolving, global outage. Datadog’s central lesson is that effective response depends less on rigid runbooks than on preparation, clear ownership, autonomous decision-making, and continuous learning. ## Datadog’s Incident Response Model - Datadog follows a “you build it, you own it” operating model. - Teams instrument their services extensively and configure monitors to detect problems around the clock. - Independent, out-of-band monitoring checks Datadog’s APIs from outside its infrastructure, ensuring that monitoring still works if Datadog itself becomes unavailable. - Slack channels are automatically created for incidents to provide shared situational awareness and enable additional engineers to contribute. ## Handling High-Severity Incidents - Senior engineers rotate on call for incidents involving substantial customer impact or multiple teams. - The first senior responder becomes the incident commander and retains overall responsibility. - A communications lead may manage internal updates and coordination. - For the most serious incidents, an engineering executive and customer-support manager join to provide leadership, business context, and customer-facing communication. - The incident commander remains accountable for coordinating the overall response. ## Preparation, Training, and Postmortems - Datadog uses a relatively low threshold for declaring incidents, giving engineers frequent practice with its response process. - Engineers complete incident-response training before joining an on-call rotation and repeat refresher training every six months. - Training covers on-call responsibilities, response roles, and blameless investigation practices. - Every high-severity incident receives a detailed postmortem focused on preventing recurrence. - Automation prompts responders to begin postmortems while the incident is still fresh. ## Autonomy and a Blameless Culture - Because large systems change constantly, detailed recovery procedures can quickly become outdated. - Datadog therefore gives engineers authority to choose the best response based on their knowledge of the affected services. - The company treats failures as weaknesses in systems rather than evidence of individual fault. - Blamelessness is intended to encourage creativity, honesty, and effective decision-making under pressure. ## The March 8 Outage - A systemd upgrade began around 06:00 UTC and ultimately triggered the outage. - Monitoring detected the problem within three minutes, and engineering teams were paged shortly afterward. - A high-severity incident was declared at 06:18, with an incident commander joining five minutes later. - The first public status update was posted at 06:31, and the outage was officially diagnosed as global at 06:32. - By 07:20, responders identified a Kubernetes failure and unhealthy intake systems as central problems. - Engineers confirmed by 08:00 that the Kubernetes failure was not spreading to additional or newly provisioned nodes. - A working mitigation for the EU1 region was found by 08:30. - Most US1 compute capacity recovered automatically by 11:00, while teams began organizing a longer recovery effort. - At 11:36, unattended upgrades were identified as the triggering event. - Compute capacity in EU1—the first step toward recovery—was restored by 12:05. ## Practical Lessons Datadog’s experience demonstrates the value of independent monitoring, practiced incident roles, rapid communication, and empowered responders. Organizations operating complex systems should regularly rehearse incident management, invest in resilient observability outside the primary platform, and use blameless postmortems to turn major outages into improvements.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into our incident response | Datadog (opens in new tab)

Datadog’s first global outage on March 8, 2023, served as a rigorous stress test for their established incident response framework and "you build it, you own it" philosophy. While the outage was triggered by a systemic failure during a routine systemd upgrade, the company's commitment to blameless culture and decentralized engineering autonomy allowed hundreds of responders to coordinate a complex recovery across multiple regions. Ultimately, the event validated their investment in out-of-band monitoring and rigorous, bi-annual incident training as essential components for managing high-scale system disasters. ## Incident Response Structure and Philosophy * Datadog employs a decentralized "you build it, you own it" model where individual engineering teams are responsible for the 24/7 health and monitoring of the services they build. * For high-severity incidents, a specialized rotation is paged, consisting of an Incident Commander to lead the response, a communications lead, and a customer liaison to manage external messaging. * The organization prioritizes "people over process," empowering engineers to use their judgment to find creative solutions rather than following rigid, pre-written playbooks that may not apply to unprecedented failures. * A blameless culture is strictly maintained across all levels of the company, ensuring that post-incident investigations focus on systemic improvements rather than assigning fault to individuals. ## Multi-Layered Monitoring Strategy * Standard telemetry provides internal visibility, but Datadog also maintains "out-of-band" monitoring that operates completely outside its own infrastructure. * This out-of-band system interacts with Datadog APIs exactly like a customer would, ensuring that engineers are alerted even if the internal monitoring platform itself becomes unavailable. * Communication is streamlined through a dedicated Slack incident app that automatically generates coordination channels, providing situational awareness to any engineer who joins the effort. ## Anatomy of the March 8 Outage * The outage began at 06:00 UTC, triggered by a systemd upgrade that caused widespread Kubernetes failures and prevented pods from restarting correctly. * The global nature of the outage was diagnosed within 32 minutes of the initial monitoring alerts, leading to the activation of executive on-calls and the customer support management team. * Responders identified "unattended upgrades" as the incident trigger approximately five and a half hours after the initial failure. * Recovery was executed in stages: compute capacity was restored first in the EU1 region, followed by the US1 region, with full infrastructure restoration completed by 19:00 UTC. Organizations should treat incident response as a perishable skill that requires constant practice through a low threshold for declaring incidents and regular training. By combining out-of-band monitoring with a culture that empowers individual engineers to act autonomously during a crisis, teams can more effectively navigate the "not if, but when" reality of large-scale system failures.

datadog2 min readCurated summary

Not just another network latency issue: How we unraveled a series of hidden bottlenecks | Datadog

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content primarily consists of Datadog’s product navigation rather than the full blog post, so it does not include Gartner’s evaluation criteria, Datadog’s specific strengths, or comparative analysis. ## Gartner Recognition - Datadog is identified as a **Leader** in Gartner’s **Magic Quadrant for Observability Platforms 2026**. - The announcement links to a Gartner-related resource hosted by Datadog. - No details are provided in the supplied text about: - Gartner’s assessment of Datadog’s execution or vision - The other vendors evaluated - Specific capabilities that led to the ranking - Any limitations or cautions noted by Gartner ## Datadog’s Observability Offering The surrounding navigation shows that Datadog positions observability as part of a broad platform covering: - Infrastructure and cloud monitoring - Application performance monitoring and profiling - Logs, metrics, databases, and data pipelines - Real user, synthetic, and mobile monitoring - Security and cloud protection - CI/CD and software delivery - Incident response, service management, and automation - AI-powered investigation and agent observability The supplied content does not explain how these products are integrated or how they support the Leader designation. Overall, the material communicates Datadog’s recognition as a Gartner observability-platform Leader, but the full article or report is needed for a meaningful technical summary.

Read original(opens in new tab)
datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level impact

Datadog’s March 8, 2023 outage was caused by an unexpected interaction between Ubuntu 22.04, systemd-networkd, and an automated security patch. A systemd change introduced behavior that flushed unfamiliar IP routing rules whenever systemd-networkd restarted; a CVE patch triggered that restart across many hosts. Because the patch was installed automatically and outside Datadog’s carefully staged deployment process, infrastructure across regions and cloud providers was affected simultaneously. ## A Systemd Behavior Change - systemd v248 introduced a systemd-networkd startup behavior that removed IP rules it did not recognize. - systemd v249 added the `ManageForeignRoutingPolicyRules` setting, which could disable this behavior, but the default configuration continued managing foreign rules. - These changes were backported to older systemd releases. - Ubuntu 20.04 used systemd v245, which did not flush IP rules during a systemd-networkd restart. - Ubuntu 22.04, adopted progressively by Datadog beginning in November 2022, used systemd v249 with the behavior enabled. Initially, the change caused no visible problems because systemd-networkd generally started only when new hosts were created, before Datadog’s custom routing rules existed. ## The Security Patch That Triggered the Problem - On March 7, 2023, Ubuntu released systemd patch `249.11-0ubuntu3.7` for a CVE. - Installing the patch restarted all systemd components, including systemd-networkd. - That restart caused systemd-networkd to flush routing policy rules on affected Ubuntu 22.04 hosts. - Ubuntu 20.04 hosts received a similar patch but were not affected because systemd v245 did not exhibit the problematic restart behavior. ## Unattended Upgrades Created Broad Exposure Datadog used Ubuntu’s default unattended-upgrade configuration: - Package metadata was downloaded twice daily using `apt-daily.timer`, with randomized delays of up to 12 hours. - Upgrades ran daily using `apt-daily-upgrade.timer`, between 06:00 and 07:00 UTC. - Only security updates and required dependencies were automatically installed. - Regular updates from the `-updates` repository were excluded. This meant many hosts automatically installed the systemd security patch during the same daily upgrade window. Not every host was affected: more than 90% of the fleet used Ubuntu 22.04, and some nodes had not yet downloaded the patch when their upgrade ran. ## Conflict with Datadog’s Deployment Process - Datadog normally updates nodes by replacing them automatically rather than relying on unattended upgrades. - Its standard process validates changes on experimental clusters, then progressively deploys them through staging and production. - Deployments are normally limited to selected clusters, availability zones, and regions before expanding. - The unattended systemd patch bypassed this process because it was installed directly on existing hosts. - As a result, a low-level networking change propagated across otherwise isolated regions and cloud providers at nearly the same time. Datadog’s experience demonstrates that even security-only automated updates can introduce coordinated infrastructure risk. Critical system packages should be tested and rolled out through the same staged process as other production changes, or their automated upgrades should be carefully constrained and monitored.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into the platform-level impact | Datadog (opens in new tab)

The March 2023 Datadog outage was triggered by a simultaneous, global failure across multiple cloud providers and regions, caused by an unexpected interaction between a systemd security patch and Ubuntu 22.04’s default networking behavior. While Datadog typically employs rigorous, staged rollouts for infrastructure changes, the automated nature of OS-level security updates bypassed these controls. The incident highlights the hidden risks in system-level defaults and the potential for "unattended upgrades" to create synchronized failures across supposedly isolated environments. ## The systemd-networkd Routing Change * In December 2020, systemd version 248 introduced a change where `systemd-networkd` flushes all IP routing rules it does not recognize upon startup. * Version 249 introduced the `ManageForeignRoutingPolicyRules` setting, which defaults to "yes," confirming this management behavior for any rules not explicitly defined in systemd configuration files. * These changes were backported to earlier versions (v247 and v248) but were notably absent from v245, the version used in Ubuntu 20.04. ## Dormant Risks in the Ubuntu 22.04 Migration * Datadog began migrating its fleet from Ubuntu 20.04 to 22.04 in late 2022, eventually reaching 90% coverage across its infrastructure. * Ubuntu 22.04 utilizes systemd v249, meaning the majority of the fleet was susceptible to the routing rule flushing behavior. * The risk remained dormant during the initial rollout because `systemd-networkd` typically only starts during the initial boot sequence when no complex routing rules have been established yet. ## The Trigger: Unattended Upgrades and the CVE Patch * On March 7, 2023, a security patch for a systemd CVE was released to the Ubuntu security repositories. * Datadog’s fleet used the Ubuntu default configuration for `unattended-upgrades`, which automatically installs security-labeled patches once a day, typically between 06:00 and 07:00 UTC. * The installation of the patch forced a restart of the `systemd-networkd` service on active, running nodes. * Upon restarting, the service identified existing IP routing rules (crucial for container networking) as "foreign" and deleted them, effectively severing network connectivity for the nodes. ## Failure of Regional Isolation * Because the security patch was released globally and the automated upgrade window was synchronized across regions, the failure occurred nearly simultaneously worldwide. * This automation bypassed Datadog’s standard practice of "baking" changes in staging and experimental clusters for weeks before proceeding to production. * Nodes on the older Ubuntu 20.04 (systemd v245) were unaffected by the patch, as that version of systemd does not flush IP rules upon a service restart. To mitigate similar risks, infrastructure teams should consider explicitly disabling the management of foreign routing rules in systemd-networkd configuration when using third-party networking plugins. Furthermore, while automated security patching is a best practice, organizations must balance the speed of patching with the need for controlled, staged rollouts to prevent global configuration drift or synchronized failures.

datadog1 min readCurated summary

Husky: Exactly-once ingestion and multi-tenancy at scale | Datadog

The provided content does not include the blog post itself. It contains Datadog’s navigation menu and a promotional banner announcing its recognition as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no article text about “Husky.” ## Content Available - The page links to Datadog products covering: - Infrastructure and application monitoring - Logs, databases, and data observability - Security and digital experience - Software delivery and service management - AI and platform capabilities - The banner promotes Datadog’s observability-platform recognition. - No technical details, sections, architecture diagrams, implementation discussion, or conclusions from the Husky article are included. Please provide the article text or a complete page extract for an accurate summary.

Read original(opens in new tab)
datadog1 min readCurated summary

DRUIDS, the design system that powers Datadog | Datadog

The provided content does not include the blog post itself; it contains Datadog navigation links and a URL titled “Druids: The Design System That Powers Datadog.” As a result, the article’s arguments, implementation details, and conclusions cannot be reliably summarized without inventing information. ## Available Information - The page appears to be a Datadog engineering blog post about **Druids**, Datadog’s design system. - The surrounding navigation lists Datadog products across: - Infrastructure and application monitoring - Logs, security, and digital experience - Software delivery and service management - AI and platform capabilities - No article sections, examples, technical architecture, or conclusions are included in the supplied text. ## Practical Conclusion Please provide the article body or a working copy of the page content for an accurate section-by-section summary.

Read original(opens in new tab)