Log Management

57 posts

datadog2 min readCurated summary

Engineering Spotlight: Tay Nishimura | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The provided content primarily consists of Datadog’s navigation and product listings, so it does not include Gartner’s evaluation details or the announcement’s supporting arguments. ## Recognition as an Observability Leader - Datadog highlights its placement as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The announcement links to a Gartner-related resource hosted by Datadog. - No specific Gartner strengths, cautions, scoring, or comparison with other vendors are included in the provided text. ## Datadog’s Broad Platform Coverage The navigation presents Datadog as a unified platform spanning: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications and data:** APM, universal service monitoring, continuous profiling, database monitoring, data-stream monitoring, and jobs monitoring. - **Logs and security:** log management, observability pipelines, sensitive-data scanning, cloud security, SIEM, workload protection, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, code coverage, incident response, SLOs, workflow automation, and case management. - **AI:** agent observability, GPU monitoring, AI integrations, AI agents, investigation tools, and an MCP server. Overall, the announcement positions Datadog’s broad, integrated observability and security platform as the basis for its Leader designation, but the supplied excerpt does not provide enough detail to assess Gartner’s underlying evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Husky, Datadog's third-generation event store | Datadog

Datadog’s page announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content does not include the report’s evaluation criteria, Gartner’s analysis, or Datadog’s supporting evidence; it primarily contains website navigation and links to Datadog products. ## Gartner Recognition - Datadog highlights its placement as a **Leader** in Gartner’s Magic Quadrant for Observability Platforms. - The announcement links to a downloadable Gartner resource. - No details are provided about Datadog’s position, strengths, weaknesses, or comparison with other vendors. ## Datadog’s Product Portfolio The page navigation presents Datadog as a broad observability and operations platform covering: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, universal service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database, data-streams, quality, jobs, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, vulnerability, compliance, SIEM, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, internal developer portals, incident response, SLOs, workflow automation, and case management. - **AI capabilities:** AI integrations, GPU monitoring, Bits AI agents, investigation tools, MCP support, and agent observability. ## What the Provided Content Does Not Cover - Gartner’s methodology or assessment criteria - Specific reasons Datadog was named a Leader - Customer feedback, market vision, or execution scores - Technical architecture, pricing, implementation guidance, or product comparisons The material supports the conclusion that Datadog is promoting broad platform coverage and third-party recognition, but the linked Gartner report would be needed for a substantive evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

How Datadog's IT team automated account inactivity and SaaS spend management | Datadog

Datadog’s IT team built an automated system to identify inactive user accounts and reduce unnecessary SaaS spending. The approach replaces manual audits with data-driven workflows that detect inactivity, notify users or owners, and reclaim unused licenses while preserving access controls and accountability. ## Automating Account Inactivity Detection - The system monitors account activity across SaaS applications. - It identifies users who have not logged in or used assigned tools for a defined period. - Automated notifications give users or managers an opportunity to confirm continued business need. - Accounts can then be suspended, deprovisioned, or escalated for review. ## Managing SaaS Spend - Inactivity data is used to find unused or underused licenses. - IT can reclaim seats instead of continuing to pay for unused subscriptions. - Usage information supports more accurate renewal and purchasing decisions. - Centralized automation reduces the manual effort required to audit many applications. ## Governance and Operational Benefits - Standardized workflows make account reviews more consistent across tools. - Automated approvals and escalation paths provide visibility into decisions. - The process helps balance cost reduction with security and user access requirements. - IT teams gain a repeatable way to manage the growing complexity of SaaS environments. Organizations with substantial SaaS usage can apply the same model: centralize activity data, define inactivity policies, automate notifications and approvals, and connect the results to license reclamation and access-management workflows.

Read original(opens in new tab)
datadog1 min readCurated summary

Profiling improvements in Go 1.18 | Datadog

The provided text does not include the blog post itself. It contains Datadog’s navigation menu and a promotional link announcing its recognition as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms, but no technical discussion or conclusions from the referenced article. ## Available content ### Datadog’s observability platform - Datadog promotes products for: - Infrastructure and Kubernetes monitoring - Application performance monitoring and profiling - Logs, databases, and data observability - Security and cloud protection - Real-user and synthetic monitoring - CI/CD and software delivery - Incident and service management - AI-powered investigation and automation ### Gartner recognition - The page links to Datadog’s announcement that it was named a Leader in the Gartner Magic Quadrant for Observability Platforms. - The excerpt does not provide Gartner’s evaluation criteria, Datadog’s strengths or weaknesses, or supporting evidence. Please provide the article body or a working text extract for an accurate technical summary.

Read original(opens in new tab)
datadog2 min readCurated summary

How Datadog's IT team automated monitoring third-party accounts | Datadog

The provided text does not contain the tech blog post itself. It consists primarily of Datadog’s navigation menu and a promotional banner announcing its “Leader” ranking in the 2026 Gartner Magic Quadrant for Observability Platforms, so the article’s argument and conclusion cannot be reliably summarized. ## Promotional Announcement - Datadog links to a Gartner report about observability platforms. - The banner presents Datadog as a Leader in the report. ## Datadog Product Categories - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, profiling, dynamic instrumentation, and agent observability. - **Data and Logs:** database monitoring, data quality, job monitoring, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, vulnerability, compliance, SIEM, workload, and application protection. - **Digital Experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, and error tracking. - **Software Delivery:** CI visibility, test optimization, code coverage, feature flags, and developer portals. - **Service Management:** event management, incident response, SLOs, workflow automation, and case management. - **AI:** AI agents, GPU monitoring, integrations, investigation tools, and MCP support. The actual blog content—apparently related to how Datadog’s IT team automated monitoring of third-party accounts—is missing. A summary would require the article body or a complete excerpt.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering spotlight: Maël Nison | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The announcement positions Datadog as a broad observability platform spanning infrastructure, applications, logs, security, digital experience, software delivery, service management, and AI. The provided content does not include Gartner’s detailed evaluation or the blog post’s supporting arguments. ## Recognition and Platform Scope - Datadog highlights its leadership placement in Gartner’s observability-platform research. - Its platform covers: - Infrastructure and container monitoring - Application performance monitoring and profiling - Database, data-stream, and jobs monitoring - Log management and observability pipelines - Cloud, application, workload, and code security - Browser and mobile real user monitoring - Synthetic monitoring, session replay, and error tracking - CI visibility, testing, code coverage, and feature flags - Incident response, service catalogs, SLOs, and workflow automation ## AI and Automation - Datadog presents AI as an integrated part of its platform through: - Bits AI agents and investigation tools - AI integrations and agent observability - GPU monitoring - MCP Server and agent-building capabilities - AI-assisted security and developer workflows - Additional automation features include Watchdog, fleet automation, workflow automation, and incident-management tools. ## Overall Positioning - The product catalog emphasizes a unified approach to monitoring technology environments rather than separate tools for infrastructure, applications, security, and user experience. - The platform also includes dashboards, alerts, notebooks, governance controls, access management, and mobile access. The announcement’s central message is that Datadog combines extensive observability coverage with security, delivery, service-management, and AI capabilities. Readers seeking the actual Gartner assessment should consult the linked Magic Quadrant resource, since the supplied text contains only the announcement and navigation information.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Glommio, a thread-per-core crate for Rust and Linux | Datadog

Datadog has been recognized as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The announcement positions Datadog as a broad observability provider spanning infrastructure, applications, logs, security, digital experience, software delivery, service management, and AI. The supplied content contains mostly site navigation rather than the article’s supporting details or Gartner’s evaluation rationale. ## Datadog’s Observability Scope - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless systems, cloud costs, GPUs, and storage. - Application performance monitoring, service monitoring, profiling, dynamic instrumentation, and agent observability. - Log management, sensitive-data scanning, audit trails, and observability pipelines. - Database, data-streams, data-quality, and jobs monitoring. ## Broader Platform Capabilities - Security features including cloud security, SIEM, workload protection, code security, vulnerability management, and compliance. - Digital-experience tools such as real-user monitoring, session replay, synthetic monitoring, error tracking, and product analytics. - Software-delivery capabilities covering CI visibility, test optimization, continuous testing, code coverage, and feature flags. - Service-management tools for incidents, events, SLOs, workflows, case management, and software catalogs. - AI offerings including Bits AI agents, investigation tools, agent observability, GPU monitoring, and MCP integrations. Overall, the announcement emphasizes Datadog’s unified and expansive observability platform. A complete assessment of Gartner’s specific strengths, cautions, and evaluation criteria would require the full blog post or linked Gartner report.

Read original(opens in new tab)
datadog2 min readCurated summary

How we wrote a Python profiler | Datadog

The post explains how Datadog built a low-overhead statistical profiler for Python. Rather than tracing every function call, the profiler periodically samples running threads and reconstructs their Python and native call stacks. The main challenge is collecting accurate stack data without pausing applications for too long or introducing unsafe behavior inside CPython. ### Why Traditional Profiling Is Expensive - Deterministic profilers instrument every function call and return. - This provides detailed data but can significantly slow production workloads. - A statistical profiler reduces overhead by sampling execution at regular intervals instead of observing every event. ### Sampling Python Threads - The profiler interrupts running threads to capture their current execution state. - Python’s signal-handling model complicates this because signals are generally processed by the main thread. - The implementation must coordinate native threads, operating-system signals, and the Python interpreter to sample worker threads reliably. - Sampling must avoid interfering with application locks or triggering unsafe operations in signal handlers. ### Reconstructing Call Stacks - A useful profile needs both Python-level frames and native stack information. - The profiler walks Python frames to identify functions, files, and line numbers. - It also handles time spent in native extensions and the interpreter itself. - Collected samples are aggregated into call stacks, allowing Datadog to show CPU usage and hotspots across the application. ### Balancing Accuracy and Overhead - Sampling frequency affects the trade-off between detail and runtime cost. - More frequent samples improve visibility into short-lived work but consume more resources. - The profiler is designed to operate continuously in production, so it prioritizes low overhead, safe memory handling, and resilience across Python versions and deployment environments. The central recommendation is to use statistical sampling for always-on production profiling. It provides actionable performance data with far less impact than call-by-call instrumentation, provided the implementation carefully accounts for CPython’s threading, signal, and native-extension behavior.

Read original(opens in new tab)
datadog4 min readCurated summary

Computing accurate percentiles with DDSketch | Datadog

Datadog’s post explains why accurately computing percentiles is difficult when monitoring large-scale, distributed systems. Traditional approaches either require retaining every observation or sacrifice accuracy through fixed-size summaries, especially for long-tailed data such as request latency. DDSketch addresses this by providing mergeable percentile estimates with a guaranteed relative-error bound and memory usage that remains effectively constant. ## Why Percentiles Matter - Averages can hide slow requests and do not describe the tail of a distribution. - Percentiles such as p95, p99, and p99.9 are more useful for measuring latency and reliability. - Monitoring systems must calculate these values from enormous numbers of observations across many hosts and services. - Storing every measurement is too expensive, while calculating percentiles independently on each machine and averaging the results is mathematically incorrect. ## Limitations of Common Approaches - Exact percentile calculation requires sorting or retaining all values, which is impractical for high-volume metrics. - Histograms use predefined buckets, making their accuracy dependent on bucket boundaries. - Fixed-width buckets are inefficient for distributions spanning several orders of magnitude: - Small values may require fine-grained buckets. - Large values may require a huge number of buckets. - Many quantile sketches optimize for rank accuracy, but a small rank error can still produce a large value error in heavy-tailed distributions. - Summaries must also be mergeable so that data collected from multiple agents can be combined without losing their accuracy guarantees. ## DDSketch’s Logarithmic Mapping - DDSketch groups values into logarithmically spaced bins rather than equally sized intervals. - Values close together near zero receive finer absolute resolution, while larger values receive wider buckets. - Each value is mapped to a key based on its logarithm: - Positive and negative values are handled separately. - Zero and values sufficiently close to zero use a dedicated zero bucket. - A representative value is chosen for each bucket, typically using the bucket’s geometric center. - Because adjacent buckets have a fixed ratio, the estimated value is bounded by a predictable relative error rather than a fixed absolute error. ## Relative-Error Guarantees - DDSketch is configured with a target relative accuracy, such as 1%. - Its logarithmic base is selected so that the returned quantile is within that relative-error bound of the true value. - Relative error is particularly appropriate for latency data: - An error of a few milliseconds matters greatly for a 10 ms request. - The same absolute error is much less significant for a 10-second request. - The sketch preserves accuracy across a wide range of values without requiring a proportional increase in the number of buckets. ## Distributed Aggregation and Memory Use - DDSketches can be merged by adding the bucket counts from separate sketches. - This allows agents, hosts, containers, and regional services to aggregate measurements into a global percentile. - Merging does not require access to the original observations. - The sketch stores counts rather than individual values, substantially reducing memory and network costs. - Datadog also describes bounded-memory variants that collapse older or less significant bins when necessary, allowing sketches to maintain a fixed storage limit while retaining useful tail information. ## Practical Trade-offs - Higher accuracy requires more buckets and therefore more memory. - Lower accuracy reduces resource usage but produces wider estimates. - The choice of relative accuracy should reflect the metric’s operational needs rather than defaulting to the smallest possible error. - Implementations must account for negative values, zeros, very small values, and values outside the normal range. - Accurate percentile reporting depends not only on the sketch algorithm but also on correct aggregation and consistent configuration across producers. DDSketch is therefore a practical choice for observability systems that need scalable, mergeable, and predictable percentile calculations. Its logarithmic buckets and relative-error guarantees make it especially well suited to latency and other long-tailed measurements where fixed-width histograms or rank-based approximations can be misleading.

Read original(opens in new tab)
datadog1 min readCurated summary

Building highly reliable data pipelines at Datadog | Datadog

The provided text does not include the blog post’s actual article content. It contains Datadog’s navigation menu and a link titled “Highly Reliable Data Pipelines,” so the post’s argument, architecture, and technical conclusions cannot be summarized reliably. ## Available Information - The page appears to be a Datadog engineering blog post about building highly reliable data pipelines. - The surrounding content is primarily Datadog product navigation. - It also promotes Datadog’s recognition as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. Please provide the article body or a complete page extract for a detailed technical summary.

Read original(opens in new tab)
datadog2 min readCurated summary

Rethinking UX for AI-driven alerting | Datadog

Datadog’s page announces that the company was named a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms. The supplied content, however, primarily contains site navigation rather than the referenced blog post, so it does not provide details about the article’s argument concerning AI-driven alerting. ## Gartner Recognition - Datadog highlights its recognition as a Leader in Gartner’s Magic Quadrant for Observability Platforms. - The announcement is presented as a promotional resource linked from the Datadog website. ## Datadog’s Product Portfolio The navigation emphasizes Datadog’s broad observability and security platform, including: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database, data-stream, job, quality, log, sensitive-data, and pipeline monitoring. - **Security:** code, cloud, vulnerability, workload, application, API, and SIEM security tools. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, and error tracking. - **Software delivery:** CI visibility, test optimization, code coverage, feature flags, and developer portals. - **Service management:** incident response, SLOs, event management, workflows, and case management. - **AI:** Bits AI agents, investigations, chat, security analysis, agent observability, and MCP integrations. The provided text does not include enough of the actual “Rethinking UX for AI-Driven Alerting” article to summarize its technical concepts or conclusions.

Read original(opens in new tab)
datadog2 min readCurated summary

Improving trust with Datadog Log Management

Datadog handles hundreds of thousands of emails daily and uses Amazon SES for critical messages such as password resets. Because SES and CloudWatch did not provide sufficiently accessible, support-friendly event data, Datadog built a serverless pipeline that forwards SES events to Datadog Log Management. This provides low-maintenance delivery infrastructure, searchable email metrics, and monitoring for failures. ## Exporting Amazon SES Events - SES configuration sets define which email events to capture: - Send - Reject - Bounce - Complaint - Delivery - Open - Click - Events are published to an Amazon SNS topic. - SNS invokes an AWS Lambda function for every event. - The Lambda forwards the event to Datadog Logs using the Datadog API key. - Terraform provisions the SNS topic, SES configuration set, event destination, IAM role, and Lambda function. - The example uses Python 2.7 and stores the API key as a Lambda environment variable; production systems should encrypt the key. - This architecture avoids maintaining a custom email service while preserving visibility into email processing. ## Making SES Events Searchable in Datadog - SES events arrive in Datadog as JSON. - Datadog’s existing AWS integration pipeline processes the logs automatically. - Important fields can be converted into facets directly from a log entry. - Datadog uses fields such as the email event type and subject to quickly search for specific password reset activity. ## Monitoring and Operational Benefits - Support teams can verify whether a recipient received or interacted with a password reset email. - The entire delivery and logging pipeline is serverless and requires minimal maintenance. - Monitors can be configured on the logs to alert normal escalation channels when any part of the pipeline fails. - The solution combines the reliability of Amazon SES with Datadog’s observability and search capabilities. Overall, routing SES events through SNS and Lambda into Datadog Log Management is a practical way to create a trusted, observable password-reset email system without operating a separate mail infrastructure.

Read original(opens in new tab)
datadog1 min readCurated summary

Improving trust with Datadog Log Management | Datadog

The provided content does not include the blog post’s main article text. It contains Datadog navigation links and a promotional banner announcing Datadog as a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms, plus a link suggesting the article concerns improving trust with Datadog Log Management. ## Datadog’s Observability Platform - Datadog promotes a broad observability platform covering: - Infrastructure and cloud monitoring - Application performance monitoring - Logs and sensitive-data protection - Security monitoring - Digital experience monitoring - Software delivery and service management - AI-powered investigation and automation - The banner highlights Datadog’s recognition as a Gartner Magic Quadrant Leader. ## Log Management and Trust - The linked article appears to focus on improving trust through Datadog Log Management. - The supplied text does not provide details about the specific problems, technologies, or recommendations discussed in the post. A complete summary requires the article’s actual body text rather than the surrounding website navigation.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Kafka-Kit: Tools for scaling Kafka | Datadog

Datadog’s “Kafka Kit” is a collection of operational tools designed to make Apache Kafka easier to scale and manage. The post argues that Kafka’s built-in administrative mechanisms become difficult to use safely as clusters grow, particularly when rebalancing partitions or adding and removing brokers. Kafka Kit automates these workflows while emphasizing balanced assignments, controlled changes, and operational visibility. ## Why Kafka Scaling Becomes Difficult - Growing Kafka clusters require frequent partition movement and broker rebalancing. - Native Kafka reassignment workflows can involve large, complex JSON configurations. - Poorly planned changes can create: - Uneven storage and traffic distribution - Excessive network and disk I/O - Overloaded brokers - Extended recovery times - Operational changes must account for replication, leadership, broker capacity, and rack or availability-zone placement. ## Kafka Kit’s Approach - Kafka Kit provides reusable tooling for common Kafka administration tasks. - The tools generate and apply partition assignments instead of requiring operators to construct them manually. - Assignments can be optimized for more even distribution of: - Partitions - Replicas - Leaders - Storage and traffic - The tooling is intended to support both routine balancing and larger cluster changes, such as adding or decommissioning brokers. ## Safer Partition Reassignment - Reassignments can be performed incrementally rather than moving all partitions at once. - Changes can be throttled to limit their effect on production workloads. - Operators can inspect proposed assignments before applying them. - Controlled movement reduces the risk of saturating Kafka brokers, disks, or network links. - The approach makes long-running migrations easier to monitor and interrupt if necessary. ## Operating Kafka at Scale - Datadog built the tools from its experience running Kafka as a critical part of its data infrastructure. - At large scale, Kafka administration needs to be repeatable and automatable rather than dependent on manual intervention. - Separating planning from execution allows teams to validate capacity and placement before changing the cluster. - Standardized tooling also helps reduce the chance of configuration errors during high-risk maintenance operations. Kafka Kit is most useful for teams operating Kafka clusters large enough that manual partition management is unreliable or disruptive. Automating assignment generation, throttling, validation, and broker lifecycle changes can make scaling more predictable and safer.

Read original(opens in new tab)