Kubernetes

144 posts

datadog3 min readCurated summary

Engineering Spotlight: Tay Nishimura

Tay Nishimura’s career shows that succeeding in tech is often less about fitting a standard engineering mold and more about finding work that matches one’s strengths. Although she initially struggled with the speed and coding demands of software development, her rigor, visual thinking, and careful approach became valuable in site reliability engineering. Her transition was enabled by self-directed learning, community education, and ToyNet, an open source networking simulator that demonstrated her practical abilities. ## Entering Tech from Mathematics - Tay began as a mathematics major focused on real analysis, then added computer science after advice from a professor. - Internships at Amazon and Google introduced her to the technology industry. - She found a sharp contrast between academia and industry: - School rewarded theoretical rigor. - Industry emphasized practical, fast, and agile solutions. - Tay also felt like an outsider because she had little exposure to computers growing up. ## Struggling with Traditional Software Engineering - Coding did not come naturally to Tay’s visual way of thinking. - She translated code into drawings to understand and modify it, then converted those ideas back into code. - This process produced high-quality, careful work but made her slower than colleagues expected. - A manager suggested product management and site reliability engineering as possible alternatives. - Tay discovered that her deliberate pace was useful for SRE work, particularly when evaluating failure modes and making critical changes. - Because her company offered no path into those roles, she eventually left rather than continue facing increasing stress. ## Discovering Networking and Technical Program Work - Tay’s next role had a software engineer title but involved work closer to product or technical program management. - She learned that job titles and responsibilities vary significantly between companies. - With better work-life balance, she began studying computer networking in her free time. - She created visual diagrams and learning modules to explain switches, routers, and packet flows. - These efforts became Project Reclass, a nonprofit teaching technical skills to incarcerated people and military veterans. - The program used improvised equipment, such as fake routers and switches, to teach concepts in environments where real networking hardware was unavailable. ## Building ToyNet During the Pandemic - After her company laid off its entire office during COVID-19, Tay decided to pursue SRE directly. - When prisons suspended in-person education, Project Reclass adapted by creating a digital networking simulator. - Tay architected ToyNet, an open source platform built with: - React - A Flask backend - Containerized Mininet instances for network emulation - Users can connect simulated routers, switches, and hosts, configure IP addresses, and run commands such as `ping` and `arp`. - ToyNet was designed to work for incarcerated learners with restricted internet access. - Deploying it in the cloud also gave Tay practical experience that helped compensate for limited professional cloud experience. - Companies interested in the project were more likely to advance her through the interview process, eventually leading to Datadog. ## Finding the Right Environment at Datadog - At Datadog, Tay learned Kubernetes, chaos engineering, network traffic control, and Go. - She found that her rigor and visual thinking were assets rather than liabilities. - While learning Datadog’s Chaos Controller codebase, she mapped files and dependencies by drawing boxes and arrows. - Her experience suggests that engineers do not need to learn or reason in a single conventional way; the right environment can turn an apparent weakness into a strength. Tay’s path recommends experimenting broadly, studying independently, and building concrete projects that reveal how you think and solve problems. The most suitable tech role may emerge only after moving between companies and disciplines rather than forcing yourself to succeed in an ill-fitting position.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering Spotlight: Tay Nishimura | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The provided content primarily consists of Datadog’s navigation and product listings, so it does not include Gartner’s evaluation details or the announcement’s supporting arguments. ## Recognition as an Observability Leader - Datadog highlights its placement as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms. - The announcement links to a Gartner-related resource hosted by Datadog. - No specific Gartner strengths, cautions, scoring, or comparison with other vendors are included in the provided text. ## Datadog’s Broad Platform Coverage The navigation presents Datadog as a unified platform spanning: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications and data:** APM, universal service monitoring, continuous profiling, database monitoring, data-stream monitoring, and jobs monitoring. - **Logs and security:** log management, observability pipelines, sensitive-data scanning, cloud security, SIEM, workload protection, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, feature flags, code coverage, incident response, SLOs, workflow automation, and case management. - **AI:** agent observability, GPU monitoring, AI integrations, AI agents, investigation tools, and an MCP server. Overall, the announcement positions Datadog’s broad, integrated observability and security platform as the basis for its Leader designation, but the supplied excerpt does not provide enough detail to assess Gartner’s underlying evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

Introducing Husky, Datadog's third-generation event store | Datadog

Datadog’s page announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The provided content does not include the report’s evaluation criteria, Gartner’s analysis, or Datadog’s supporting evidence; it primarily contains website navigation and links to Datadog products. ## Gartner Recognition - Datadog highlights its placement as a **Leader** in Gartner’s Magic Quadrant for Observability Platforms. - The announcement links to a downloadable Gartner resource. - No details are provided about Datadog’s position, strengths, weaknesses, or comparison with other vendors. ## Datadog’s Product Portfolio The page navigation presents Datadog as a broad observability and operations platform covering: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, universal service monitoring, continuous profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database, data-streams, quality, jobs, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, workload, vulnerability, compliance, SIEM, and application/API protection. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, experiments, and error tracking. - **Software delivery and service management:** CI visibility, test optimization, internal developer portals, incident response, SLOs, workflow automation, and case management. - **AI capabilities:** AI integrations, GPU monitoring, Bits AI agents, investigation tools, MCP support, and agent observability. ## What the Provided Content Does Not Cover - Gartner’s methodology or assessment criteria - Specific reasons Datadog was named a Leader - Customer feedback, market vision, or execution scores - Technical architecture, pricing, implementation guidance, or product comparisons The material supports the conclusion that Datadog is promoting broad platform coverage and third-party recognition, but the linked Gartner report would be needed for a substantive evaluation.

Read original(opens in new tab)
datadog2 min readCurated summary

How Datadog's IT team automated account inactivity and SaaS spend management | Datadog

Datadog’s IT team built an automated system to identify inactive user accounts and reduce unnecessary SaaS spending. The approach replaces manual audits with data-driven workflows that detect inactivity, notify users or owners, and reclaim unused licenses while preserving access controls and accountability. ## Automating Account Inactivity Detection - The system monitors account activity across SaaS applications. - It identifies users who have not logged in or used assigned tools for a defined period. - Automated notifications give users or managers an opportunity to confirm continued business need. - Accounts can then be suspended, deprovisioned, or escalated for review. ## Managing SaaS Spend - Inactivity data is used to find unused or underused licenses. - IT can reclaim seats instead of continuing to pay for unused subscriptions. - Usage information supports more accurate renewal and purchasing decisions. - Centralized automation reduces the manual effort required to audit many applications. ## Governance and Operational Benefits - Standardized workflows make account reviews more consistent across tools. - Automated approvals and escalation paths provide visibility into decisions. - The process helps balance cost reduction with security and user access requirements. - IT teams gain a repeatable way to manage the growing complexity of SaaS environments. Organizations with substantial SaaS usage can apply the same model: centralize activity data, define inactivity policies, automate notifications and approvals, and connect the results to license reclamation and access-management workflows.

Read original(opens in new tab)
datadog3 min readCurated summary

It's always DNS . . . except when it's not: A deep dive through gRPC, Kubernetes, and AWS networking

A routine update to a critical metrics query service caused intermittent errors and increased latency. Although logs initially pointed to DNS failures, the investigation revealed a deeper networking problem involving dropped packets and saturated AWS VPC connection tracking. The incident highlighted how Kubernetes, Cilium, AWS networking, and DNS behavior can interact in ways that obscure the true cause. ## Initial Symptoms and Apparent DNS Failures - Errors increased whenever the metrics query service was rolled out. - The service retrieves data from metric stores for dashboards and monitor evaluations. - Automatic retries reduced user-facing failures but increased latency. - Service logs showed DNS errors when connecting to dependencies inside Kubernetes. - The investigation therefore began with the cluster’s DNS infrastructure. ## NodeLocal DNSCache Reaches Its Limits - NodeLocal DNSCache runs as a `node-local-dns` DaemonSet on every Kubernetes node. - DNS pods had: - A 64 MB memory limit - A `max_concurrent` limit of 1,000 requests - The pods experienced out-of-memory errors and rejected requests during rollouts. - Increasing memory to 256 MB stopped the OOM errors, but DNS failures continued. - Request volume was far below the expected capacity: - Normally about 400 queries per second - Nearly 2,000 queries per second during rollouts - Expected capacity of at least 200,000 queries per second - Upstream resolvers were marked unhealthy, suggesting that NodeLocal DNSCache could not establish or maintain connections. - Because upstream requests could wait up to five seconds, connection failures consumed concurrency slots and made the cache appear overloaded. ## Evidence of a Network Problem - The instances were below their 5-Gbps sustained throughput limits. - TCP retransmits increased in correlation with service rollouts. - Engineers suspected brief traffic spikes, or microbursts, that were not visible in aggregate throughput metrics. - This shifted the investigation from DNS configuration toward lower-level AWS networking behavior. ## AWS VPC Connection Tracking - ENA metrics revealed a significant increase in `conntrack_allowance_exceeded`. - This metric counts packets dropped when VPC connection tracking becomes saturated. - Connection tracking maintains state for network flows and supports features such as stateful EC2 security groups. - The infrastructure used two tracking layers: - VPC conntrack maintained at the hypervisor level - Linux conntrack inside each instance - VPC conntrack appeared saturated even though Linux conntrack contained fewer than 60,000 entries—well within the observed capacity of similar instances. - AWS Support confirmed that conntrack capacity varies by instance type and that VPC conntrack limits could differ substantially from Linux conntrack limits. - Scaling to larger instances resolved the symptoms, but the engineers wanted to understand the traffic pattern and find a more efficient long-term solution. ## VPC Flow Logs as the Next Investigation Tool - The team turned to Amazon VPC Flow Logs to examine the service’s low-level network behavior. - These logs were expected to clarify why connection tracking filled up and how rollout traffic contributed to the saturation. - The investigation was still ongoing at the point where the provided article excerpt ends.

Read original(opens in new tab)
datadogOriginal article

Using the Dirty Pipe vulnerability to break out from containers | Datadog (opens in new tab)

The Dirty Pipe vulnerability (CVE-2022-0847) is a critical Linux kernel flaw that allows unprivileged processes to write data to any file they can read, effectively bypassing standard write permissions. This primitive is particularly dangerous in containerized environments like Kubernetes, where it can be leveraged to overwrite the host’s container runtime binary. By exploiting how the kernel manages page caches, an attacker can achieve a full container breakout and gain administrative privileges on the underlying host. ## Container Runtimes and the OCI Specification * Kubernetes utilizes the Container Runtime Interface (CRI) to manage containers via high-level runtimes like containerd or CRI-O. * These high-level runtimes rely on low-level Open Container Interface (OCI) runtimes, most commonly runC, to handle the heavy lifting of namespaces and control groups. * Isolation is achieved by runC setting up a restricted environment before executing the user-supplied entrypoint via the `execve` system call. ## Evolution of runC Vulnerabilities * A historical vulnerability, CVE-2019-5736, previously allowed escapes by overwriting the host’s runC binary through the `/proc/self/exe` file descriptor. * To mitigate this, runC was updated to either clone the binary before execution or mount the host's runC binary as read-only inside the container. * While the read-only mount improved performance through kernel cache page sharing, it created a target for the Dirty Pipe vulnerability, which specifically targets the kernel page cache. ## The Dirty Pipe Exploitation Primitive * Dirty Pipe allows an attacker to overwrite any file they can read, including read-only files, by manipulating the kernel's internal pipe-buffer structures. * The exploit targets the page cache, meaning the overwrite is non-persistent and resides only in memory; the original file on disk remains unchanged. * In a container escape scenario, the attacker waits for a runC process to start (triggered by actions like `kubectl exec`) and targets the file descriptor at `/proc/<runC-pid>/exe`. ## Proof-of-Concept Escape Walkthrough * The attack begins with a standard, unprivileged pod running a malicious script that monitors the system for new runC processes. * Once a `kubectl exec` command is issued by an administrator, the script identifies the runC PID and applies the Dirty Pipe exploit to the associated executable. * The exploit overwrites the runC binary in the kernel page cache with a malicious ELF binary. * Because the host kernel is executing this hijacked binary with root privileges to manage the container, the attacker’s malicious code (e.g., a reverse shell or administrative command) runs with full host-level authority. To protect against this attack vector, it is essential to patch the Linux kernel to a version that includes the fix for CVE-2022-0847 and ensure that container nodes are running updated distributions.

datadog3 min readCurated summary

Escaping containers using the Dirty Pipe vulnerability | Datadog Security Labs

The post demonstrates how the Linux Dirty Pipe vulnerability can enable an unprivileged process to escape a container and gain administrative privileges on the host. The exploit abuses runC’s execution model and its host binary, which is exposed read-only inside the container but can still be modified through the kernel page cache. A proof of concept shows how a compromised Kubernetes pod can overwrite runC with a malicious executable when an administrator runs `kubectl exec`. ## Container Runtimes and runC - Kubernetes commonly uses containerd or CRI-O through the Container Runtime Interface (CRI). - These runtimes rely on lower-level OCI runtimes, most notably runC, to create isolated Linux processes. - runC configures namespaces, cgroups, and the container environment before executing the supplied entrypoint with `execve`. - During execution, `/proc/self/exe` inside the container can refer to an open descriptor for the runC binary on the host. ## Earlier runC Escape Vulnerability - CVE-2019-5736 exploited this `/proc/self/exe` behavior: - A malicious container entrypoint could write to the host’s runC binary. - Overwriting runC enabled subsequent container operations to execute attacker-controlled code with host-level privileges. - runC initially mitigated the issue by cloning its binary before execution. - It later changed the design to mount the runC binary read-only inside the container, improving performance through kernel page-cache sharing. - That optimization created conditions in which Dirty Pipe could bypass the apparent read-only protection. ## Dirty Pipe as a Container Escape Primitive - Dirty Pipe allows an unprivileged process to overwrite files it can read, even without write permission. - The modification occurs in the kernel page cache rather than persistent storage: - The original file remains intact on disk. - Dropping caches or rebooting can restore the original contents. - Despite being temporary, the overwrite is sufficient to execute malicious code when the modified binary is run. - In this case, the attacker targets the host’s runC binary through `/proc/<runC-pid>/exe`. ## Kubernetes Proof of Concept - The demonstration starts an ordinary, unprivileged pod using an attacker-controlled container image. - Its entrypoint script: - Replaces `/bin/sh` with a launcher referencing `/proc/self/exe`. - Waits for a runC process to appear. - Invokes the Dirty Pipe exploit against that process’s executable. - An administrator running `kubectl exec` causes runC to execute inside the container, triggering the overwrite. - The modified runC is replaced with a malicious ELF binary that runs commands such as `id` and `hostname`, recording their output in `/tmp/hacked`. - The exploit is adapted from the original Dirty Pipe proof of concept and the earlier runC escape technique. The attack illustrates that kernel vulnerabilities can undermine container isolation even when the container is unprivileged and the target binary is mounted read-only. Systems should promptly patch vulnerable Linux kernels and container runtimes, while treating compromised containers as potential paths to host compromise.

Read original(opens in new tab)
datadog1 min readCurated summary

Profiling improvements in Go 1.18 | Datadog

The provided text does not include the blog post itself. It contains Datadog’s navigation menu and a promotional link announcing its recognition as a Leader in Gartner’s 2026 Magic Quadrant for Observability Platforms, but no technical discussion or conclusions from the referenced article. ## Available content ### Datadog’s observability platform - Datadog promotes products for: - Infrastructure and Kubernetes monitoring - Application performance monitoring and profiling - Logs, databases, and data observability - Security and cloud protection - Real-user and synthetic monitoring - CI/CD and software delivery - Incident and service management - AI-powered investigation and automation ### Gartner recognition - The page links to Datadog’s announcement that it was named a Leader in the Gartner Magic Quadrant for Observability Platforms. - The excerpt does not provide Gartner’s evaluation criteria, Datadog’s strengths or weaknesses, or supporting evidence. Please provide the article body or a working text extract for an accurate technical summary.

Read original(opens in new tab)
datadogOriginal article

Our journey taking Kubernetes state metrics to the next level | Datadog (opens in new tab)

Datadog’s container observability team significantly improved the performance of kube-state-metrics (KSM) by contributing core architectural enhancements to the upstream open-source project. Faced with scalability bottlenecks where metrics collection for large clusters took tens of seconds and generated massive data payloads, they revamped the underlying library to achieve a 15x improvement in processing duration. These contributions allowed for high-granularity monitoring at scale, ensuring that the Datadog Agent can efficiently handle millions of metrics across thousands of Kubernetes nodes. ### Challenges with KSM Scalability * KSM uses the informer pattern to expose cluster-level metadata via the Openmetrics format, but the volume of data grows exponentially with cluster size. * In high-scale environments, a single node generates approximately nine metrics, while a single pod can generate up to 40 metrics. * In clusters with thousands of nodes and tens of thousands of pods, the `/metrics` endpoint produced payloads weighing tens of megabytes. * The time required to crawl these metrics often exceeded 15 seconds, forcing administrators to reduce check frequency and sacrifice real-time data granularity. ### Limitations of Legacy Implementations * KSM v1 relied on a monolithic loop that instantiated a Builder to track resources via stores, but it lacked efficient hooks for metric generation. * The original Python-based Datadog Agent check struggled with the "data dump" approach of KSM, where all metrics were processed at once during query time. * To manage the load, Datadog was forced to split KSM into multiple deployments based on resource types (e.g., separate deployments for pods, nodes, and secondary resources like services or deployments). * This fragmentation made the infrastructure more complex to manage and did not solve the fundamental issue of inefficient metric serialization. ### Architectural Improvements in KSM v2.0 * Datadog collaborated with the upstream community during the development of KSM v2.0 to introduce a more extensible design. * The team focused on improving the Builder and metric generation hooks to prevent the system from dumping the entire dataset at query time. * By moving away from the restrictive v1 library structure, they enabled more efficient reconciliation of metric names and metadata joins. * The resulting 15x performance gain allows the Datadog Agent to reconcile labels and tags—such as joining deployment labels to specific metrics—without the significant latency overhead previously experienced. Contributing back to the open-source community proved more effective than maintaining internal forks for scaling Kubernetes infrastructure. Organizations running high-density clusters should prioritize upgrading to KSM v2.0 and optimizing their agent configurations to leverage these architectural improvements for better observability performance.

datadog3 min readCurated summary

Our journey taking Kubernetes state metrics to the next level

Datadog contributed major scalability improvements to kube-state-metrics (KSM), after discovering that its metric collection process struggled with very large Kubernetes clusters. Millions of metrics could require tens of megabytes and tens of seconds to process every 15 seconds, forcing Datadog to reduce collection frequency. Their redesign improved collection duration by 15x and enabled more granular monitoring at scale. ## Datadog’s Kubernetes Observability Role - The Datadog Containers team monitors Kubernetes infrastructure and ensures reliable collection of: - Logs - Traces - Custom metrics - Profiles - Security signals - KSM is central to Datadog products such as Kubernetes metrics integration and Orchestrator Explorer. ## How Kubernetes State Metrics Works - KSM uses Kubernetes informers to watch objects registered with the API server. - Enabled collectors monitor resources such as pods, nodes, deployments, and services. - It generates lifecycle and metadata metrics in text-based OpenMetrics format. - Users can restrict monitored resources through the `resources` flag. - The Datadog Agent’s KSM check: - Runs every 15 seconds. - Discovers KSM containers. - Crawls their `/metrics` endpoint. - Reconciles metric metadata and applies configured label joins. - Label joins allow metadata from one metric, such as a deployment label, to become a tag on other metrics for the same object. ## Scaling Challenges - Datadog found that KSM needed to be split across multiple deployments beyond a few hundred nodes and thousands of pods. - Their deployments divided collectors by resource type: - Pods - Nodes - Services, deployments, jobs, persistent volumes, and other resources - Metric volume varied substantially: - Endpoints, jobs, and deployments produced roughly five metrics per object. - Nodes produced around nine metrics each. - Pods produced around 40 metrics each. - Large clusters with thousands of nodes and tens of thousands of pods could generate millions of metrics per scrape. - Crawling the metrics endpoint could take tens of seconds and transfer tens of megabytes. - Datadog had to reduce check frequency, sacrificing metric granularity and user experience. ## KSM’s Original Architecture - KSM v1 relied on a central loop that created a Builder and managed resource stores. - Each store used informers to track a particular Kubernetes resource. - For example, an HPA store maintained the list-and-watch logic for HorizontalPodAutoscalers. - The Builder generated metrics from the tracked resources. - Datadog identified two limitations: - Too much data was emitted and processed at query time. - The Builder did not provide a suitable extension point for custom metric-generation logic. ## Contributing the Redesign Upstream - As the KSM community prepared version 2.0 in early 2020, Datadog saw an opportunity to address its scalability and extensibility problems in the upstream project. - Rather than maintaining a private solution, the team contributed its findings and improvements to the open-source community. - The resulting work reportedly reduced metric collection duration by 15x, making high-scale, more frequent collection practical. Datadog’s experience shows that upstream open-source collaboration can solve internal infrastructure bottlenecks while improving the project for the broader Kubernetes community.

Read original(opens in new tab)
datadog2 min readCurated summary

How Datadog's IT team automated monitoring third-party accounts | Datadog

The provided text does not contain the tech blog post itself. It consists primarily of Datadog’s navigation menu and a promotional banner announcing its “Leader” ranking in the 2026 Gartner Magic Quadrant for Observability Platforms, so the article’s argument and conclusion cannot be reliably summarized. ## Promotional Announcement - Datadog links to a Gartner report about observability platforms. - The banner presents Datadog as a Leader in the report. ## Datadog Product Categories - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, profiling, dynamic instrumentation, and agent observability. - **Data and Logs:** database monitoring, data quality, job monitoring, log management, sensitive-data scanning, and observability pipelines. - **Security:** code, cloud, vulnerability, compliance, SIEM, workload, and application protection. - **Digital Experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, and error tracking. - **Software Delivery:** CI visibility, test optimization, code coverage, feature flags, and developer portals. - **Service Management:** event management, incident response, SLOs, workflow automation, and case management. - **AI:** AI agents, GPU monitoring, integrations, investigation tools, and MCP support. The actual blog content—apparently related to how Datadog’s IT team automated monitoring of third-party accounts—is missing. A summary would require the article body or a complete excerpt.

Read original(opens in new tab)
datadog1 min readCurated summary

How we minimized the overhead of Kubernetes in our job system | Datadog

The supplied content does not include the blog post itself; it contains Datadog’s navigation menu and a link titled “Moving a Job System to Kubernetes.” As a result, there is not enough article text to accurately summarize its arguments, implementation details, or conclusions. ## Available Information - The linked post appears to concern migrating a job-processing system to Kubernetes. - The surrounding page lists Datadog products for: - Infrastructure and Kubernetes monitoring - Application performance monitoring - Logs, databases, and jobs - Security and software delivery - No technical discussion, architecture description, challenges, or results from the post is included. Please provide the article’s body or a working page extract for a detailed summary.

Read original(opens in new tab)
datadog3 min readCurated summary

How we minimized the overhead of Kubernetes in our job system

Kubernetes can improve machine management and scalability, but its scheduling and runtime overhead can significantly reduce job throughput if configured poorly. Datadog found its Kubernetes-based job system used more CPU and completed jobs 40–50% more slowly than the previous VM-based system. By designing a controlled experiment, choosing better metrics, and tuning pod resource requests, the team recovered performance to roughly VM parity while investigating the overhead of running one parent process per pod. ## Designing a Comparable Experiment - The initial comparison was difficult because the Kubernetes and VM deployments differed in: - Number of nodes - Number of worker-parent clusters - Workload and enqueue rate - A controlled experiment was created using: - Identical `c5.2xlarge` machines - The same kernel version, `3.13.0-141` - Both systems repeatedly running a simple Python job - Each Kubernetes pod contained one parent process and its worker processes, making pod count equivalent to parent-process count per node. - The older kernel did not include CPU mitigations, avoiding that variable in the comparison. ## Choosing Useful Performance Metrics ### Measuring node effort - Load average initially appeared useful for measuring machine utilization. - Kubernetes background processes—such as cluster polling and pod-state checks—artificially increased load average. - Load average counts runnable processes rather than the amount of CPU time they actually consume. - The team therefore used CPU idle time instead: - It measures unused CPU capacity. - It reflects actual CPU work rather than the number of active processes. ### Measuring system performance - The job system optimized for throughput rather than latency. - Throughput was measured by the number of jobs completed within 30 seconds. - Latency remained useful for detecting queueing problems, but throughput was the primary success metric. ## Tuning Kubernetes Resource Requests - The main performance gains came from improving pod scheduling. - The target was six pods per `c5.2xlarge` node. - Initially, each pod requested: - One full CPU core - More memory than necessary - Since the node had eight cores and approximately 1.5 GiB of memory consumed by Kubernetes and system services, only four pods could be scheduled. - Requests were reduced to: - `100m` CPU, or 100 millicores - `500 MB` memory - CPU tuning generally enabled six pods per node, although some nodes still scheduled only five. - Further memory reduction was needed because system daemons consumed enough memory to prevent six pods from fitting on some nodes. - Resource requests affect scheduling minimums, while limits constrain containers after they start. - These request changes did not slow jobs because the pods still received sufficient resources to operate. ## One Parent Process per Pod - The team considered placing multiple parent processes in each pod to reduce potential pod overhead. - One parent plus its workers was a natural application unit and simplified orchestration. - The decision depended on how much overhead each pod introduced: - High overhead would favor fewer, larger pods. - Low overhead would favor one parent per pod for simpler management. - Using `pstree`, the team identified six job-system instances per node and traced their process trees through components such as: - `containerd-shim` - `tini` - The application process - They estimated that each pod included overhead associated with three containers, particularly `containerd-shim`. - CPU overhead was then investigated using `perf sched`. The practical lesson is to compare equivalent workloads, measure actual CPU consumption rather than relying blindly on load average, and tune Kubernetes requests for the desired packing density. Resource requests should be large enough for reliable operation but not so large that they unnecessarily prevent pods from being scheduled together.

Read original(opens in new tab)
datadog2 min readCurated summary

Engineering spotlight: Maël Nison | Datadog

Datadog announces that it has been named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. The announcement positions Datadog as a broad observability platform spanning infrastructure, applications, logs, security, digital experience, software delivery, service management, and AI. The provided content does not include Gartner’s detailed evaluation or the blog post’s supporting arguments. ## Recognition and Platform Scope - Datadog highlights its leadership placement in Gartner’s observability-platform research. - Its platform covers: - Infrastructure and container monitoring - Application performance monitoring and profiling - Database, data-stream, and jobs monitoring - Log management and observability pipelines - Cloud, application, workload, and code security - Browser and mobile real user monitoring - Synthetic monitoring, session replay, and error tracking - CI visibility, testing, code coverage, and feature flags - Incident response, service catalogs, SLOs, and workflow automation ## AI and Automation - Datadog presents AI as an integrated part of its platform through: - Bits AI agents and investigation tools - AI integrations and agent observability - GPU monitoring - MCP Server and agent-building capabilities - AI-assisted security and developer workflows - Additional automation features include Watchdog, fleet automation, workflow automation, and incident-management tools. ## Overall Positioning - The product catalog emphasizes a unified approach to monitoring technology environments rather than separate tools for infrastructure, applications, security, and user experience. - The platform also includes dashboards, alerts, notebooks, governance controls, access management, and mobile access. The announcement’s central message is that Datadog combines extensive observability coverage with security, delivery, service-management, and AI capabilities. Readers seeking the actual Gartner assessment should consult the linked Magic Quadrant resource, since the supplied text contains only the announcement and navigation information.

Read original(opens in new tab)
datadog1 min readCurated summary

PHP 8: Observability baked right in | Datadog

Datadog announces that Gartner named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The supplied content contains the announcement headline and Datadog’s product navigation, but not the report’s evaluation details, methodology, or supporting arguments. ## Gartner Recognition - Datadog is positioned as a Leader in Gartner’s Magic Quadrant for Observability Platforms. - The linked resource appears to provide the full Gartner report or announcement. - No specific Gartner strengths, cautions, rankings, or comparison with other vendors are included in the provided text. ## Datadog’s Observability Portfolio The page navigation highlights Datadog’s broad platform, including: - Infrastructure monitoring, metrics, containers, Kubernetes, networks, serverless, and cloud costs - APM, service monitoring, profiling, and dynamic instrumentation - Database, data-stream, job, and quality monitoring - Log management, observability pipelines, and sensitive-data scanning - Real-user monitoring, session replay, synthetic monitoring, and error tracking - CI visibility, testing, code coverage, and software delivery tools - Incident response, service catalogs, SLOs, workflow automation, and event management - AI capabilities such as agent observability, Bits AI, GPU monitoring, and MCP integrations Overall, the material presents Datadog as a broad, integrated observability platform, but the actual Gartner analysis is not included. For a detailed assessment, consult the linked Gartner report directly.

Read original(opens in new tab)