Observability

106 posts

datadog2 min readCurated summary

.NET Continuous Profiler: Exception and lock contention | Datadog

Datadog announces that it has been named a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. The provided content, however, contains only the announcement link and Datadog’s website navigation; it does not include the underlying technical blog post or its arguments. ## Datadog’s Observability Offering - The site organizes products across: - Infrastructure monitoring, containers, Kubernetes, networks, serverless, and cloud costs - APM, continuous profiling, dynamic instrumentation, and agent observability - Database, data-stream, job, and quality monitoring - Logs, sensitive-data scanning, audit trails, and observability pipelines - Security, cloud security, SIEM, workload protection, and code security - Digital experience monitoring, session replay, synthetic monitoring, and error tracking - CI visibility, test optimization, code coverage, feature flags, and developer tools - Incident response, service catalogs, SLOs, workflow automation, and case management - AI agents, GPU monitoring, AI integrations, and MCP tooling ## Missing Blog Content - The URL path references “.NET Continuous Profiler – Part 3,” but the supplied excerpt does not contain that article’s text. - No profiling techniques, implementation details, performance findings, or conclusions are provided. - A meaningful technical summary would require the full blog post content. The available material supports only the conclusion that Datadog is promoting its recognition as a Gartner observability-platform Leader and positioning its broad product portfolio as part of that platform.

Read original(opens in new tab)
datadog1 min readCurated summary

Engineering spotlight: Marie-Laure Bardonnet | Datadog

The provided content does not include the actual blog post. It contains Datadog’s navigation menu and a promotional banner announcing its recognition as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no article text to summarize. ## Datadog’s Observability Announcement - Datadog promotes its recognition as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked announcement appears to be a marketing resource rather than the blog post itself. ## Datadog Product Areas Listed - Infrastructure monitoring, metrics, containers, Kubernetes, networking, serverless, and cloud costs - Application performance monitoring, profiling, dynamic instrumentation, and database monitoring - Log management, observability pipelines, and sensitive data scanning - Security, including cloud security, SIEM, workload protection, and code security - Digital experience tools such as RUM, session replay, synthetic monitoring, and error tracking - Software delivery, CI visibility, testing, feature flags, and code coverage - Service management, incident response, workflows, dashboards, and AI capabilities The article body or source text is needed for a meaningful section-by-section summary.

Read original(opens in new tab)
datadog2 min readCurated summary

.NET Continuous Profiler: CPU and wall time profiling | Datadog

Datadog’s page announces that the company was named a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. The provided content mainly consists of Datadog’s product navigation, showing the breadth of its observability, security, digital experience, software delivery, and AI offerings; it does not include the blog article’s substantive text. ## Gartner Recognition - Datadog highlights its position as a Leader in the **Gartner Magic Quadrant for Observability Platforms 2026**. - The page links to a resource describing this recognition. ## Datadog’s Platform Coverage - **Infrastructure:** infrastructure, container, network, serverless, cloud cost, storage, and GPU monitoring. - **Applications:** APM, universal service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database monitoring, data-stream monitoring, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, application, API, SIEM, vulnerability, compliance, and entitlement management. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, incident response, SLOs, workflow automation, and case management. - **AI capabilities:** AI agents, investigation tools, GPU monitoring, integrations, MCP services, and AI-assisted development. The supplied excerpt supports the conclusion that Datadog is presenting Gartner’s recognition as validation of its broad, integrated observability platform. A detailed technical summary would require the full blog post, which is not included here.

Read original(opens in new tab)
datadog2 min readCurated summary

.NET Continuous Profiler: Under the hood | Datadog

Datadog is presented as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The page highlights Datadog’s broad platform, spanning infrastructure, applications, data, logs, security, digital experience, software delivery, service management, and AI. However, the supplied content contains mostly navigation links rather than the blog post’s substantive analysis. ## Gartner Recognition - Datadog’s featured announcement is its designation as a **Leader** in the Gartner® Magic Quadrant™ for Observability Platforms. - The linked resource appears to provide the full Gartner-related announcement and evaluation details. ## Broad Observability Platform The listed Datadog capabilities cover: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, universal service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database monitoring, data-stream monitoring, job and quality monitoring, log management, and observability pipelines. - **Security:** cloud security, SIEM, vulnerability management, code security, workload protection, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, incident response, SLOs, workflow automation, and case management. - **AI:** GPU monitoring, AI integrations, Bits AI agents, investigation tools, and an MCP server. ## Overall Takeaway The available material positions Datadog as a unified observability and operations platform with capabilities extending well beyond traditional infrastructure monitoring. For the Gartner evaluation criteria, supporting evidence, and detailed rationale behind the Leader designation, the full linked article or report would be required.

Read original(opens in new tab)
datadog1 min readCurated summary

Scaling Self-Serve Analytics: The Tools Empowering 5,000 Employees | Datadog

The provided content does not include the blog post itself. It contains Datadog’s navigation menu and a promotional banner announcing its position as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no technical discussion from the article. ## Available Content - The page appears to be a Datadog engineering blog post related to a CrunchConf talk on self-serve analytics. - The supplied text primarily lists Datadog products across: - Infrastructure and application monitoring - Data and log management - Security - Digital experience - Software delivery - Service management - AI capabilities - A separate banner promotes Datadog’s recognition by Gartner. ## Missing Article Details - No article title, introduction, body sections, technical examples, architecture, or conclusions are included. - The available text is insufficient to accurately summarize the post’s arguments or implementation details. Please provide the blog post’s main text or a complete extraction of the page for a substantive summary.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering spotlight: Jeromy Carriere | Datadog

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content does not include the report’s evaluation criteria or detailed rationale, but it presents Datadog as a broad observability platform spanning infrastructure, applications, data, logs, security, digital experience, software delivery, and AI. ## Gartner Recognition - Datadog highlights its position as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The linked resource appears to provide the full Gartner-related announcement and assessment. ## Broad Observability Coverage - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database monitoring, data-stream and job monitoring, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, vulnerability, compliance, SIEM, and application/API protection. - **Digital experience:** real-user monitoring, session replay, synthetic monitoring, product analytics, mobile testing, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, incident response, SLOs, workflow automation, and developer portals. - **AI capabilities:** AI agents, investigation tools, GPU monitoring, integrations, MCP support, and AI-assisted chat and coding. Datadog’s positioning rests on consolidating telemetry and operational workflows across the technology stack. To understand the Gartner recognition in depth, readers would need the linked report, since the supplied excerpt contains the announcement and product navigation but not the supporting analysis.

Read original(opens in new tab)
datadog3 min readCurated summary

Engineering spotlight: Jeromy Carriere

Jeromy Carriere, Datadog’s SVP of Product Engineering, describes engineering leadership as balancing strategy, execution, people development, and organizational processes. His career across Google, Facebook, and Datadog shaped his passion for observability and taught him to lead through mistakes, autonomy, and accountability. He argues that sustained innovation requires both intentional direction and space for teams and individuals to grow. ## The Responsibilities of Engineering Leadership - Carriere’s work shifts with organizational cycles: - Quarterly planning focused on connecting initiatives and increasing collaboration. - Execution periods focused on removing resource and decision-making blockers. - Ongoing performance management across the broader engineering organization. - He works on how Engineering operates, including: - Process definition and improvement. - Hiring and performance management. - Reviewing design documents and code. - Observing incidents and postmortems. - His central challenge is balancing attention across strategy, execution, people, and technical quality. ## From Cloud Monitoring to Datadog - At Google in 2014, Carriere helped create a cloud monitoring offering because Google Cloud lacked capabilities comparable to Datadog. - He later worked on observability at Facebook and developed a strong interest in improving developer and engineer productivity. - He returned to Datadog after seeing its ability to innovate and deliver products with sustained velocity. - He emphasizes that velocity means more than moving quickly: it requires direction, strategy, and consistency over time. ## Learning Through Mistakes - Carriere believes the most valuable lessons come from making and owning mistakes. - Earlier in his career, he was sometimes too directive, limiting team creativity and ownership. - At other times, he was too distant and failed to provide enough support. - His current leadership approach aims to: - Give teams substantial autonomy. - Provide support when needed. - Hold teams accountable for agreed-upon outcomes. - He also learned that people may not have a clear five-year career plan. Leaders should help them identify the work that provides satisfaction and enables them to perform at their best. ## Creating Space for Career Decisions - People often become focused on the immediate task and overlook other possibilities. - Carriere recommends deliberately stepping back to observe: - What activities feel satisfying. - What opportunities exist nearby. - What kinds of work could better match an individual’s strengths and interests. - This reflection requires intentional time and freedom rather than waiting for clarity to emerge automatically. ## The Value of Co-op and Internship Programs - Carriere credits the University of Waterloo’s co-op program with giving him early experience as a professional software developer. - The combination of strong academic training and repeated, high-quality industry placements helped connect theory with real work. - He sees a similar benefit in Datadog’s internship program, where interns are trusted with meaningful projects and often produce some of the company’s strongest work. Datadog’s engineering approach, as described by Carriere, combines strategic product velocity with thoughtful organizational support. For both leaders and individual contributors, the practical recommendation is to learn from mistakes, create room for reflection, and build environments where people have autonomy, meaningful work, and accountability.

Read original(opens in new tab)
datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level recovery

Datadog’s March 8, 2023 outage removed 60% of its compute capacity, forcing teams to restore infrastructure in stages while accounting for regional and cloud-provider differences. In EU1, recovery depended on rebooting affected nodes, restoring Kubernetes control planes in a strict hierarchy, and gradually bringing application capacity back online. Scaling afterward exposed infrastructure limits that had not been considered during normal operations. ## EU1 Platform Recovery - A system patch disconnected affected EU1 nodes from the network, but the nodes could be recovered through reboots. - Recovery was initially slowed by the lack of observability and unavailable Kubernetes APIs. - Datadog operates: - **Parent clusters**, which host the control-plane pods for other clusters. - **Child clusters**, where Datadog applications run. - This hierarchy allows Datadog to use Kubernetes deployment, replacement, rolling-update, and autoscaling capabilities for child-cluster control planes. - Parent-cluster control planes run on VMs and are managed with `systemd`. ## Restoring Kubernetes Clusters Because both parent and child environments were affected by the Ubuntu 22.04 issue, recovery had to follow a strict sequence: - **Parent control planes:** Nodes running Cilium were rebooted to restore network connectivity. This finished by 08:45 UTC. - **Child control planes:** All parent-cluster nodes hosting child control-plane pods were rebooted. This finished by 09:30 UTC. - **Application nodes:** Thousands of instances across dozens of child clusters were restarted. - Recovery reached 60% by 10:20 UTC. - All application nodes were restored by 12:05 UTC. - Restarts were prioritized by workload importance and paced to avoid overwhelming Kubernetes control planes. ## Scaling Capacity and Recovering Backlogs After restoring the clusters, Datadog needed substantial additional capacity to process data buffered during the outage. - EU1 hit a Google Cloud mesh limit of **15,500 VM instances** at 14:18 UTC. - Instance creation failures became apparent around 15:00 UTC. - Datadog had not checked this documented limit before the incident, but Google Cloud quickly raised it after Datadog submitted a high-priority request. - Autoscaling also exhausted the IP capacity of subnets used by three log- and trace-processing clusters. - These clusters normally used about 35–45% of their IP capacity, but the backlog caused autoscaling to request more than twice their usual replica counts, filling the subnets. ## Practical Lessons The recovery demonstrated that restoring compute capacity is not enough: teams must also understand dependency order, control-plane architecture, cloud-provider quotas, and network-address limits. Capacity planning should account for severe backlog-driven scaling, not just normal operating utilization, and documented infrastructure limits should be validated before emergencies occur.

Read original(opens in new tab)
datadog2 min readCurated summary

Not just another network latency issue: How we unraveled a series of hidden bottlenecks | Datadog

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content primarily consists of Datadog’s product navigation rather than the full blog post, so it does not include Gartner’s evaluation criteria, Datadog’s specific strengths, or comparative analysis. ## Gartner Recognition - Datadog is identified as a **Leader** in Gartner’s **Magic Quadrant for Observability Platforms 2026**. - The announcement links to a Gartner-related resource hosted by Datadog. - No details are provided in the supplied text about: - Gartner’s assessment of Datadog’s execution or vision - The other vendors evaluated - Specific capabilities that led to the ranking - Any limitations or cautions noted by Gartner ## Datadog’s Observability Offering The surrounding navigation shows that Datadog positions observability as part of a broad platform covering: - Infrastructure and cloud monitoring - Application performance monitoring and profiling - Logs, metrics, databases, and data pipelines - Real user, synthetic, and mobile monitoring - Security and cloud protection - CI/CD and software delivery - Incident response, service management, and automation - AI-powered investigation and agent observability The supplied content does not explain how these products are integrated or how they support the Leader designation. Overall, the material communicates Datadog’s recognition as a Gartner observability-platform Leader, but the full article or report is needed for a meaningful technical summary.

Read original(opens in new tab)
datadog1 min readCurated summary

Husky: Exactly-once ingestion and multi-tenancy at scale | Datadog

The provided content does not include the blog post itself. It contains Datadog’s navigation menu and a promotional banner announcing its recognition as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, but no article text about “Husky.” ## Content Available - The page links to Datadog products covering: - Infrastructure and application monitoring - Logs, databases, and data observability - Security and digital experience - Software delivery and service management - AI and platform capabilities - The banner promotes Datadog’s observability-platform recognition. - No technical details, sections, architecture diagrams, implementation discussion, or conclusions from the Husky article are included. Please provide the article text or a complete page extract for an accurate summary.

Read original(opens in new tab)
datadog1 min readCurated summary

DRUIDS, the design system that powers Datadog | Datadog

The provided content does not include the blog post itself; it contains Datadog navigation links and a URL titled “Druids: The Design System That Powers Datadog.” As a result, the article’s arguments, implementation details, and conclusions cannot be reliably summarized without inventing information. ## Available Information - The page appears to be a Datadog engineering blog post about **Druids**, Datadog’s design system. - The surrounding navigation lists Datadog products across: - Infrastructure and application monitoring - Logs, security, and digital experience - Software delivery and service management - AI and platform capabilities - No article sections, examples, technical architecture, or conclusions are included in the supplied text. ## Practical Conclusion Please provide the article body or a working copy of the page content for an accurate section-by-section summary.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering Spotlight: Tay Nishimura | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The provided content primarily consists of Datadog’s navigation and product listings, so it does not include Gartner’s evaluation details or the announcement’s supporting arguments. ## Recognition as an Observability Leader - Datadog highlights its placement as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The announcement links to a Gartner-related resource hosted by Datadog. - No specific Gartner strengths, cautions, scoring, or comparison with other vendors are included in the provided text. ## Datadog’s Broad Platform Coverage The navigation presents Datadog as a unified platform spanning: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications and data:** APM, universal service monitoring, continuous profiling, database monitoring, data-stream monitoring, and jobs monitoring. - **Logs and security:** log management, observability pipelines, sensitive-data scanning, cloud security, SIEM, workload protection, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, code coverage, incident response, SLOs, workflow automation, and case management. - **AI:** agent observability, GPU monitoring, AI integrations, AI agents, investigation tools, and an MCP server. Overall, the announcement positions Datadog’s broad, integrated observability and security platform as the basis for its Leader designation, but the supplied excerpt does not provide enough detail to assess Gartner’s underlying evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Husky, Datadog's third-generation event store | Datadog

Datadog’s page announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content does not include the report’s evaluation criteria, Gartner’s analysis, or Datadog’s supporting evidence; it primarily contains website navigation and links to Datadog products. ## Gartner Recognition - Datadog highlights its placement as a **Leader** in Gartner’s Magic Quadrant for Observability Platforms. - The announcement links to a downloadable Gartner resource. - No details are provided about Datadog’s position, strengths, weaknesses, or comparison with other vendors. ## Datadog’s Product Portfolio The page navigation presents Datadog as a broad observability and operations platform covering: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, universal service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database, data-streams, quality, jobs, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, vulnerability, compliance, SIEM, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, internal developer portals, incident response, SLOs, workflow automation, and case management. - **AI capabilities:** AI integrations, GPU monitoring, Bits AI agents, investigation tools, MCP support, and agent observability. ## What the Provided Content Does Not Cover - Gartner’s methodology or assessment criteria - Specific reasons Datadog was named a Leader - Customer feedback, market vision, or execution scores - Technical architecture, pricing, implementation guidance, or product comparisons The material supports the conclusion that Datadog is promoting broad platform coverage and third-party recognition, but the linked Gartner report would be needed for a substantive evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

How Datadog's IT team automated account inactivity and SaaS spend management | Datadog

Datadog’s IT team built an automated system to identify inactive user accounts and reduce unnecessary SaaS spending. The approach replaces manual audits with data-driven workflows that detect inactivity, notify users or owners, and reclaim unused licenses while preserving access controls and accountability. ## Automating Account Inactivity Detection - The system monitors account activity across SaaS applications. - It identifies users who have not logged in or used assigned tools for a defined period. - Automated notifications give users or managers an opportunity to confirm continued business need. - Accounts can then be suspended, deprovisioned, or escalated for review. ## Managing SaaS Spend - Inactivity data is used to find unused or underused licenses. - IT can reclaim seats instead of continuing to pay for unused subscriptions. - Usage information supports more accurate renewal and purchasing decisions. - Centralized automation reduces the manual effort required to audit many applications. ## Governance and Operational Benefits - Standardized workflows make account reviews more consistent across tools. - Automated approvals and escalation paths provide visibility into decisions. - The process helps balance cost reduction with security and user access requirements. - IT teams gain a repeatable way to manage the growing complexity of SaaS environments. Organizations with substantial SaaS usage can apply the same model: centralize activity data, define inactivity policies, automate notifications and approvals, and connect the results to license reclamation and access-management workflows.

Read original(opens in new tab)
datadog1 min readCurated summary

It's always DNS . . . except when it's not: A deep dive through gRPC, Kubernetes, and AWS networking | Datadog

The supplied content does not include the blog post itself; it mainly contains Datadog navigation links and a promotional banner stating that Datadog was named a Leader in Gartner’s Magic Quadrant for Observability Platforms. The URL suggests the missing article concerns a gRPC, DNS, and load-balancing incident, but no incident details are provided. ## Available Content - Datadog promotes its recognition as a Gartner Magic Quadrant Leader. - The navigation lists products across: - Infrastructure and application monitoring - Logs, databases, and data observability - Security - Digital experience monitoring - Software delivery and service management - AI and platform capabilities - The linked page path references an engineering post titled “gRPC, DNS, and Load Balancing Incident.” ## Missing Technical Details - No description of the incident or its impact - No explanation of the DNS or load-balancing failure - No timeline, root-cause analysis, or remediation steps - No lessons learned or recommendations Please provide the full article text for a substantive technical summary.

Read original(opens in new tab)