Cloud Computing

30 posts

awsOriginal article

Amazon EC2 X8i instances powered by custom Intel Xeon 6 processors are generally available for memory-intensive workloads (opens in new tab)

Amazon has announced the general availability of EC2 X8i instances, specifically engineered for memory-intensive workloads such as SAP HANA, large-scale databases, and data analytics. Powered by custom Intel Xeon 6 processors with a 3.9 GHz all-core turbo frequency, these instances provide a significant performance leap over the previous X2i generation. By offering up to 6 TB of memory and substantial improvements in throughput, X8i instances represent the highest-performing Intel-based memory-optimized option in the AWS cloud. ### Performance Enhancements and Processor Architecture * **Custom Silicon:** The instances utilize custom Intel Xeon 6 processors available exclusively on AWS, delivering the fastest memory bandwidth among comparable Intel cloud processors. * **Memory and Bandwidth:** X8i provides 1.5 times more memory capacity (up to 6 TB) and 3.4 times more memory bandwidth compared to previous-generation X2i instances. * **Workload Benchmarks:** Real-world performance gains include a 50% increase in SAP Application Performance Standard (SAPS), 47% faster PostgreSQL performance, 88% faster Memcached performance, and a 46% boost in AI inference. ### Scalable Instance Sizes and Throughput * **Flexible Sizing:** The instances are available in 14 sizes, including new larger formats such as the 48xlarge, 64xlarge, and 96xlarge. * **Bare Metal Options:** Two bare metal sizes (metal-48xl and metal-96xl) are available for workloads requiring direct access to physical hardware resources. * **Networking and Storage:** The architecture supports up to 100 Gbps of network bandwidth with Elastic Fabric Adapter (EFA) support and up to 80 Gbps of Amazon EBS throughput. * **Bandwidth Control:** Support for Instance Bandwidth Configuration (IBC) allows users to customize the allocation of performance between networking and EBS to suit specific application needs. ### Cost Efficiency and Use Cases * **Licensing Optimization:** In preview testing, customers like Orion reduced SQL Server licensing costs by 50% by maintaining performance thresholds with fewer active cores compared to older instance types. * **Enterprise Applications:** The instances are SAP-certified, making them ideal for RISE with SAP and other high-demand ERP environments. * **Broad Utility:** Beyond databases, the instances are optimized for Electronic Design Automation (EDA) and complex data analytics that require massive memory footprints. For organizations managing massive datasets or expensive licensed database software, migrating to X8i instances offers a clear path to both performance optimization and infrastructure cost reduction. These instances are currently available in the US East (N. Virginia), US West (Oregon), and Europe (Ireland) regions through On-Demand, Spot, and Reserved purchasing models.

awsOriginal article

AWS Weekly Roundup: AWS Lambda for .NET 10, AWS Client VPN quickstart, Best of AWS re:Invent, and more (January 12, 2026) (opens in new tab)

The AWS Weekly Roundup for January 2026 highlights a significant push toward modernization, headlined by the introduction of .NET 10 support for AWS Lambda and Apache Airflow 2.11 for Amazon MWAA. To encourage exploration of these and other emerging technologies, AWS has revamped its Free Tier to offer new users up to $200 in credits and six months of risk-free experimentation. These updates collectively aim to streamline serverless development, enhance container storage efficiency, and provide more robust authentication options for messaging services. ### Modernized Runtimes and Orchestration * AWS Lambda now supports .NET 10 as both a managed runtime and a container base image, with AWS providing automatic updates to these environments as they become available. * Amazon Managed Workflows for Apache Airflow (MWAA) has added support for version 2.11, which serves as a critical stepping stone for users preparing to migrate to Apache Airflow 3. ### Infrastructure and Resource Management * Amazon ECS has extended support for `tmpfs` mounts to Linux tasks running on AWS Fargate and Managed Instances; this allows developers to utilize memory-backed file systems for containerized workloads to avoid writing sensitive or temporary data to task storage. * AWS Config has expanded its monitoring capabilities to discover, assess, and audit new resource types across Amazon EC2, Amazon SageMaker, and Amazon S3 Tables. * A new AWS Client VPN quickstart was released, providing a CloudFormation template and a step-by-step guide to automate the deployment of secure client-to-site VPN connections. ### Security and Messaging Enhancements * Amazon MQ for RabbitMQ brokers now supports HTTP-based authentication, which can be enabled and managed through the broker’s configuration file. * RabbitMQ brokers on Amazon MQ also now support certificate-based authentication using mutual TLS (mTLS) to improve the security posture of messaging applications. ### Educational Initiatives and Community Events * New AWS Free Tier accounts now include a 6-month trial period featuring $200 in credits and access to over 30 always-free services, specifically targeting developers interested in AI/ML and compute experimentation. * AWS published a curated "Best of re:Invent 2025" playlist, featuring high-impact sessions and keynotes for those who missed the live event. * The 2026 AWS Summit season begins shortly, with upcoming events scheduled for Dubai on February 10 and Paris on March 10. Developers should take immediate advantage of the new .NET 10 Lambda runtime for serverless applications and review the updated ECS `tmpfs` documentation to optimize container performance. For those new to the platform, the expanded Free Tier credits provide an excellent opportunity to prototype AI/ML workloads with minimal financial risk.

awsOriginal article

AWS Weekly Roundup: Amazon ECS, Amazon CloudWatch, Amazon Cognito and more (December 15, 2025) (opens in new tab)

The AWS Weekly Roundup for mid-December 2025 highlights a series of updates designed to streamline developer workflows and enhance security across the cloud ecosystem. Following the momentum of re:Invent 2025, these releases focus on reducing operational friction through faster database provisioning, more granular container control, and AI-assisted development tools. These advancements collectively aim to simplify infrastructure management while providing deeper cost visibility and improved performance for enterprise applications. ## Database and Developer Productivity * **Amazon Aurora DSQL** now supports near-instant cluster creation, reducing provisioning time from minutes to seconds to facilitate rapid prototyping and AI-powered development via the Model Context Protocol (MCP) server. * **Amazon Aurora PostgreSQL** has integrated with **Kiro powers**, allowing developers to use AI-assisted coding for schema management and database queries through pre-packaged MCP servers. * **Amazon CloudWatch SDK** introduced support for optimized JSON and CBOR protocols, improving the efficiency of data transmission and processing within the monitoring suite. * **Amazon Cognito** simplified user communications by enabling automated email delivery through Amazon SES using verified identities, removing the need for manual SES configuration. ## Compute and Networking Optimizations * **Amazon ECS on AWS Fargate** now honors custom container stop signals, such as SIGQUIT or SIGINT, allowing for graceful shutdowns of applications that do not use the default SIGTERM instruction. * **Application Load Balancer (ALB)** received performance enhancements that reduce latency for establishing new connections and lower resource consumption during traffic processing. * **AWS Fargate** cost optimization strategies were highlighted in new technical guides, focusing on leveraging Graviton processors and Fargate Spot to maximize compute efficiency. ## Security and Cost Management * **Amazon WorkSpaces Secure Browser** introduced Web Content Filtering, providing category-based access control across 25+ predefined categories and granular URL policies at no additional cost. * **AWS Cost Management** tools now feature **Tag Inheritance**, which automatically applies tags from resources to cost data, allowing for more precise tracking in Cost Explorer and AWS Budgets. * **Amazon Step Functions** integration with Amazon Bedrock was further detailed in community resources, showcasing how to build resilient, long-running AI workflows with integrated error handling. To take full advantage of these updates, organizations should review their Fargate task definitions to implement custom stop signals for better application stability and enable Tag Inheritance to improve the accuracy of year-end cloud financial reporting.

awsOriginal article

AWS Weekly Roundup: AWS re:Invent keynote recap, on-demand videos, and more (December 8, 2025) (opens in new tab)

The December 8, 2025, AWS Weekly Roundup recaps the major themes from AWS re:Invent, signaling a significant industry transition from AI assistants to autonomous AI agents. While technical innovation in infrastructure remains a priority, the event underscored that developers remain at the heart of the AWS mission, empowered by new tools to automate complex tasks using natural language. This shift represents a "renaissance" in cloud computing, where purpose-built infrastructure is now designed to support the non-deterministic nature of agentic workloads. ## Community Recognition and the Now Go Build Award * Raphael Francis Quisumbing (Rafi) from the Philippines was honored with the Now Go Build Award, presented by Werner Vogels. * A veteran of the ecosystem, Quisumbing has served as an AWS Hero since 2015 and has co-led the AWS User Group Philippines for over a decade. * The recognition emphasizes AWS's continued focus on community dedication and the role of individual builders in empowering regional developer ecosystems. ## The Evolution from AI Assistants to Agents * AWS CEO Matt Garman identified AI agents as the next major inflection point for the industry, moving beyond simple chat interfaces to systems that perform tasks and automate workflows. * Dr. Swami Sivasubramanian highlighted a paradigm shift where natural language serves as the primary interface for describing complex goals. * These agents are designed to autonomously generate plans, write necessary code, and call various tools to execute complete solutions without constant human intervention. * AWS is prioritizing the development of production-ready infrastructure that is secure and scalable specifically to handle the "non-deterministic" behavior of these AI agents. ## Core Infrastructure and the Developer Renaissance * Despite the focus on AI, AWS reaffirmed that its core mission remains the "freedom to invent," keeping developers central to its 20-year strategy. * Leaders Peter DeSantis and Dave Brown reinforced that foundational attributes—security, availability, and performance—remain the non-negotiable pillars of the AWS cloud. * The integration of AI agents is framed as a way to finally realize material business returns on AI investments by moving from experimental use cases to automated business logic. To maximize the value of these updates, organizations should begin evaluating how to transition from simple LLM implementations to agentic frameworks that can execute end-to-end business processes. Reviewing the on-demand keynote sessions from re:Invent 2025 is recommended for technical teams looking to implement the latest secure, agent-ready infrastructure.

awsOriginal article

Introducing checkpointless and elastic training on Amazon SageMaker HyperPod (opens in new tab)

Amazon SageMaker HyperPod has introduced checkpointless and elastic training features to accelerate AI model development by minimizing infrastructure-related downtime. These advancements replace traditional, slow checkpoint-restart cycles with peer-to-peer state recovery and enable training workloads to scale dynamically based on available compute capacity. By decoupling training progress from static hardware configurations, organizations can significantly reduce model time-to-market while maximizing cluster utilization. **Checkpointless Training and Rapid State Recovery** * Replaces the traditional five-stage recovery process—including job termination, network setup, and checkpoint retrieval—which can often take up to an hour on self-managed clusters. * Utilizes peer-to-peer state replication and in-process recovery to allow healthy nodes to restore the model state instantly without restarting the entire job. * Incorporates technical optimizations such as collective communications initialization and memory-mapped data loading to enable efficient data caching. * Reduces recovery downtime by over 80% based on internal studies of clusters with up to 2,000 GPUs, and was a core technology used in the development of Amazon Nova models. **Elastic Training and Automated Cluster Scaling** * Allows AI workloads to automatically expand to use idle cluster capacity as it becomes available and contract when resources are needed for higher-priority tasks. * Reduces the need for manual intervention, saving hours of engineering time previously spent reconfiguring training jobs to match fluctuating compute availability. * Optimizes total cost of ownership by ensuring that training momentum continues even as inference volumes peak and pull resources away from the training pool. * Orchestrates these transitions seamlessly through the HyperPod training operator, ensuring that model development is not disrupted by infrastructure changes. For teams managing large-scale AI workloads, adopting these features can reclaim significant development time and lower operational costs by preventing idle cluster periods. Organizations scaling to thousands of accelerators should prioritize checkpointless training to mitigate the impact of hardware faults and maintain continuous training momentum.

googleOriginal article

Solving virtual machine puzzles: How AI is optimizing cloud computing (opens in new tab)

Google researchers have developed LAVA, a scheduling framework designed to optimize virtual machine (VM) allocation in large-scale data centers by accurately predicting and adapting to VM lifespans. By moving beyond static, one-time predictions toward a "continuous re-prediction" model based on survival analysis, the system significantly improves resource efficiency and reduces fragmentation. This approach allows cloud providers to solve the complex "bin packing" problem more effectively, leading to better capacity utilization and easier system maintenance. ### The Challenge of Long-Tailed VM Distributions * Cloud workloads exhibit a extreme long-tailed distribution: while 88% of VMs live for less than an hour, these short-lived jobs consume only 2% of total resources. * The rare VMs that run for 30 days or longer account for a massive fraction of compute resources, meaning their placement has a disproportionate impact on host availability. * Poor allocation leads to "resource stranding," where a server's remaining capacity is too small or unbalanced to host new VMs, effectively wasting expensive hardware. * Traditional machine learning models that provide only a single prediction at VM creation are often fragile, as a single misprediction can block a physical host from being cleared for maintenance or new tasks. ### Continuous Re-prediction via Survival Analysis * Instead of predicting a single average lifetime, LAVA uses an ML model to generate a probability distribution of a VM's expected duration. * The system employs "continuous re-prediction," asking how much longer a VM is expected to run given how long it has already survived (e.g., a VM that has run for five days is assigned a different remaining lifespan than a brand-new one). * This adaptive approach allows the scheduling logic to automatically correct for initial mispredictions as more data about the VM's actual behavior becomes available over time. ### Novel Scheduling and Rescheduling Algorithms * **Non-Invasive Lifetime Aware Scheduling (NILAS):** Currently deployed on Google’s Borg cluster manager, this algorithm ranks potential hosts by grouping VMs with similar expected exit times to increase the frequency of "empty hosts" available for maintenance. * **Lifetime-Aware VM Allocation (LAVA):** This algorithm fills resource gaps on hosts containing long-lived VMs with jobs that are at least an order of magnitude shorter. This ensures the short-lived VMs exit quickly without extending the host's overall occupation time. * **Lifetime-Aware Rescheduling (LARS):** To minimize disruptions during defragmentation, LARS identifies and migrates the longest-lived VMs first while allowing short-lived VMs to finish their tasks naturally on the original host. By integrating survival-analysis-based predictions into the core logic of data center management, cloud providers can transition from reactive scheduling to a proactive model. This system not only maximizes resource density but also ensures that the physical infrastructure remains flexible enough to handle large, resource-intensive provisioning requests and essential system updates.

datadog2 min readCurated summary

Unraveling a Postgres segfault that uncovered an Arm64 JIT compiler bug

Postgres was crashing with segmentation faults when executing certain expensive queries on an Arm64 Kubernetes cluster. Investigators reduced the failure to a simple table scan and discovered that disabling JIT compilation prevented the crash. Assembly-level debugging ultimately traced the problem to a bug in LLVM’s Arm64 JIT support. ## Isolating the Crash - The failures occurred across multiple EC2 nodes, ruling out faulty hardware. - Query logs showed that the crashes consistently followed a small number of query patterns. - The simplest reproducer was: ```sql SELECT repo_id FROM repository; ``` - Core dumps had badly corrupted stacks, but surviving frames pointed to `ExecRunCompiledExpr`, suggesting a failure during JIT execution. - The unusually short backtraces reinforced the suspicion that the stack itself had been corrupted. ## How PostgreSQL JIT Works - PostgreSQL normally evaluates SQL expressions through a general-purpose interpreter. - JIT compilation converts expressions such as `1+1` into native machine code, reducing interpreter overhead for large workloads. - JIT can also optimize tuple deforming by converting disk tuples into in-memory values more efficiently. - PostgreSQL uses LLVM to generate the compiled code. - Because compilation adds overhead and compiled functions are not reused between queries, PostgreSQL enables JIT primarily for expensive queries based on cost thresholds. ## The Query of Death - The affected query scanned a partitioned `repository` table with: - 64 partitions - More than 1.6 million rows - 128 JIT-generated functions - Its query plan enabled expression compilation and tuple deforming: ```text JIT: Functions: 128 Expressions: true Deforming: true Inlining: false Optimization: false ``` - Running the query with: ```sql SET jit = off; ``` completed successfully. - Disabling JIT cluster-wide immediately stopped the crashes without noticeable query-latency effects. ## Root Cause Direction - The release build of PostgreSQL offered limited debugging flexibility, so the team planned to reproduce the failure in a dedicated test environment. - Further investigation eventually isolated the defect to JIT compilation on Arm64 systems. - The underlying issue was identified as an LLVM bug rather than a PostgreSQL query or hardware problem. - The investigation continued down to generated assembly and resulted in an upstream fix. The immediate mitigation was to disable PostgreSQL JIT, while the durable solution was to adopt the LLVM fix addressing the Arm64 code-generation bug.

Read original(opens in new tab)
datadog3 min readCurated summary

How we built a Ruby library that saves 50% in testing time

Datadog built a Ruby test impact analysis library to reduce CI time and avoid rerunning unrelated, flaky tests. The approach maps each test to the source files it executes, then runs only tests affected by a commit. Existing Ruby coverage APIs were too slow or incompatible with standard coverage tools, so Datadog developed a faster solution using Ruby VM interpreter events. ## The CI Problem - Large test suites often take 20 minutes or more and may fail because of unrelated flaky tests. - Parallel execution reduces runtime but increases cloud costs and does not eliminate flakiness. - Selective testing can reduce: - Pipeline duration - Cloud resource usage - Exposure to unrelated flaky tests - Test impact analysis determines which source files each test executes and compares them with files changed in the latest Git commit. ## Requirements for Test Impact Analysis Datadog’s library needed to provide: - **Correctness:** Never skip a test that could detect a regression. - **Performance:** Add minimal overhead because impact data must be collected on every commit and branch. - **Seamlessness:** Require no user code changes and avoid changing test behavior or interfering with existing tooling. ## Limitations of Ruby Coverage APIs - Ruby’s built-in `Coverage` module can collect per-test coverage using `resume` and `suspend`, introduced in Ruby 3.1. - A prototype using Coverage had two major problems: - It conflicted with tools such as SimpleCov that collect total code coverage. - It added up to 300% overhead, making the test suite roughly four times slower. - Datadog then tried Ruby’s `TracePoint` API, subscribing to the `line` VM event. - TracePoint avoided interference with SimpleCov and provided the required data, but still introduced roughly 200% overhead, reaching 400% in some cases. ## A Custom Coverage Tool Using Ruby VM Events - Datadog examined Ruby’s internals, including: - `coverage.c` - `rb_coverage_resume` - `rb_resume_coverages` - `rb_add_event_hook2` - Ruby’s C extension API supports registering callbacks for `RUBY_EVENT_LINE`, allowing a custom native implementation. - The proof of concept: - Registers a line-event hook for the current thread. - Records the source file for executed lines. - Ignores files outside the project root. - Removes the hook when collection stops. - Returns and resets the collected coverage data for the next test. - Implementing collection closer to the VM was intended to preserve correctness while substantially reducing the overhead of per-test impact tracking. Datadog’s experience shows that selective testing is a promising way to make CI faster and more reliable, but practical test impact analysis requires a low-level implementation. Standard coverage and tracing APIs provide useful functionality but can impose unacceptable performance costs, making a purpose-built native VM-event collector a better fit.

Read original(opens in new tab)
datadogOriginal article

2023-03-08 incident: A deep dive into the platform-level recovery | Datadog (opens in new tab)

Following a massive system-wide outage in March 2023, Datadog successfully restored its EU1 region by identifying that a simple node reboot could resolve network connectivity issues caused by a faulty system patch. While the team managed to restore 100 percent of compute capacity within hours, the recovery effort was subsequently hindered by cloud provider infrastructure limits and IP address exhaustion. This post-mortem highlights the complexities of scaling hierarchical Kubernetes environments under extreme pressure and the importance of accounting for "black swan" capacity requirements. ## Hierarchical Kubernetes Recovery Datadog utilizes a strict hierarchy of Kubernetes clusters to manage its infrastructure, which necessitated a granular, three-tiered recovery approach. Because the outage affected network connectivity via `systemd-networkd`, the team had to restore components in a specific order to regain control of the environment. * **Parent Control Planes:** Engineers first rebooted the virtual machines hosting the parent clusters, which manage the control planes for all other clusters. * **Child Control Planes:** Once parent clusters were stable, the team restored the control planes for application clusters, which run as pods within the parent infrastructure. * **Application Worker Nodes:** Thousands of worker nodes across dozens of clusters were restarted progressively to avoid overwhelming the control planes, reaching full capacity by 12:05 UTC. ## Scaling Bottlenecks and Cloud Quotas Once the infrastructure was online, the team attempted to scale out rapidly to process a massive backlog of buffered data. This surge in demand triggered previously unencountered limitations within the Google Cloud environment. * **VPC Peering Limits:** At 14:18 UTC, the platform hit a documented but overlooked limit of 15,500 VM instances within a single network peering group, blocking all further scaling. * **Provider Intervention:** Datadog worked directly with Google Cloud support to manually raise the peering group limit, which allowed scaling to resume after a nearly four-hour delay. ## IP Address and Subnet Capacity Even after cloud-level instance quotas were lifted, specific high-traffic clusters processing logs and traces hit a secondary bottleneck related to internal networking. * **Subnet Exhaustion:** These clusters attempted to scale to more than twice their normal size, quickly exhausting all available IP addresses in their assigned subnets. * **Capacity Planning Gaps:** While Datadog typically targets a 66% maximum IP usage to allow for a 50% scale-out, the extreme demands of the recovery backlog exceeded these safety margins. * **Impact on Backlog:** For six hours, the lack of available IPs forced these clusters to process data significantly slower than the rest of the recovered infrastructure. ## Recovery Summary The EU1 recovery demonstrates that even when hardware is functional, software-defined limits can create cascading delays. Organizations should not only monitor their own resource usage but also maintain visibility into cloud provider quotas and ensure that subnet allocations account for extreme recovery scenarios where workloads may need to double or triple in size momentarily.

datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level recovery

Datadog’s March 8, 2023 outage removed 60% of its compute capacity, forcing teams to restore infrastructure in stages while accounting for regional and cloud-provider differences. In EU1, recovery depended on rebooting affected nodes, restoring Kubernetes control planes in a strict hierarchy, and gradually bringing application capacity back online. Scaling afterward exposed infrastructure limits that had not been considered during normal operations. ## EU1 Platform Recovery - A system patch disconnected affected EU1 nodes from the network, but the nodes could be recovered through reboots. - Recovery was initially slowed by the lack of observability and unavailable Kubernetes APIs. - Datadog operates: - **Parent clusters**, which host the control-plane pods for other clusters. - **Child clusters**, where Datadog applications run. - This hierarchy allows Datadog to use Kubernetes deployment, replacement, rolling-update, and autoscaling capabilities for child-cluster control planes. - Parent-cluster control planes run on VMs and are managed with `systemd`. ## Restoring Kubernetes Clusters Because both parent and child environments were affected by the Ubuntu 22.04 issue, recovery had to follow a strict sequence: - **Parent control planes:** Nodes running Cilium were rebooted to restore network connectivity. This finished by 08:45 UTC. - **Child control planes:** All parent-cluster nodes hosting child control-plane pods were rebooted. This finished by 09:30 UTC. - **Application nodes:** Thousands of instances across dozens of child clusters were restarted. - Recovery reached 60% by 10:20 UTC. - All application nodes were restored by 12:05 UTC. - Restarts were prioritized by workload importance and paced to avoid overwhelming Kubernetes control planes. ## Scaling Capacity and Recovering Backlogs After restoring the clusters, Datadog needed substantial additional capacity to process data buffered during the outage. - EU1 hit a Google Cloud mesh limit of **15,500 VM instances** at 14:18 UTC. - Instance creation failures became apparent around 15:00 UTC. - Datadog had not checked this documented limit before the incident, but Google Cloud quickly raised it after Datadog submitted a high-priority request. - Autoscaling also exhausted the IP capacity of subnets used by three log- and trace-processing clusters. - These clusters normally used about 35–45% of their IP capacity, but the backlog caused autoscaling to request more than twice their usual replica counts, filling the subnets. ## Practical Lessons The recovery demonstrated that restoring compute capacity is not enough: teams must also understand dependency order, control-plane architecture, cloud-provider quotas, and network-address limits. Capacity planning should account for severe backlog-driven scaling, not just normal operating utilization, and documented infrastructure limits should be validated before emergencies occur.

Read original(opens in new tab)
figma2 min readCurated summary

Designing in the cloud with confidence | Figma Blog

Figma’s post announces that it has received the EU Cloud Code of Conduct compliance mark, making it one of only 17 companies to earn the designation at the time. The certification signals that Figma’s privacy and security practices have been independently validated against GDPR-related cloud data protection standards. Figma presents this as reassurance for global teams, particularly organizations operating in or expanding into Europe. ## Commitment to Global, Secure Collaboration - Figma says scaling design work across regions requires a reliable and trustworthy product. - Its international strategy includes: - Building features for users around the world - Expanding its global presence - Protecting customers’ design intellectual property - Meeting regional security and compliance requirements - The company emphasizes transparency around its security practices and track record. ## EU Cloud Code of Conduct Compliance - Figma became one of only 17 companies to receive the EU Cloud Code of Conduct compliance mark. - The Code is described as a leading European standard for cloud data protection. - The mark indicates that Figma has established processes intended to handle user data in compliance with the EU’s General Data Protection Regulation (GDPR). - Figma’s privacy and security policies were validated by SCOPE Europe, the accredited monitoring body for the Code. ## What It Means for Customers - Figma users can have greater confidence that their design data is handled according to recognized security standards. - European organizations and companies planning international expansion can use Figma with clearer assurance about its compliance posture. - The company directs readers to the EU Cloud Code of Conduct and its broader security documentation for more information. Figma’s practical message is that customers can continue using its cloud-based design platform with increased confidence in its European data protection and security controls.

Read original(opens in new tab)
figma3 min readCurated summary

Figma and Chromebook: Empowering the next generation of designers | Figma Blog

Figma’s partnership with Google for Education brings Figma and FigJam to Chromebooks, making professional design and collaboration tools available to millions of students at no cost to participating schools. The initiative aims to remove financial, hardware, and geographic barriers to design education while preparing students for a digital-first workforce. Figma argues that design skills—especially visual communication, collaboration, problem-solving, and creativity—should be accessible to everyone, not only students with specialized degrees or expensive equipment. ## Bringing Design Tools to Chromebooks - School districts can deploy and manage free Figma Organization licenses through the Google Admin Console. - Figma and FigJam will be available on Chromebooks, the most common personal computing devices for students. - Students can design, brainstorm, and collaborate from their own school devices rather than relying on limited computer-lab sessions. - Schools can apply for the beta program beginning in summer 2022. ## Removing Barriers to Design Education - Specialized degrees, costly hardware, and expensive creative software can prevent students from gaining design experience. - Cloud-based Chromebooks make design tools more affordable and accessible. - Figma’s founders designed the product to be free because they had seen how expensive creative tools could be for college students. - The broader goal is to give students access to opportunities in design and software regardless of their financial background. ## Why Design Skills Matter - The shift from physical to digital work, accelerated by the pandemic, has increased the importance of digital design and visual communication. - Design education develops: - Complex problem-solving - Collaboration - Creative expression - Communication and critical thinking - These skills are valuable across disciplines, not only in traditional design careers. - Figma cites research suggesting that design-focused companies outperform their peers by a factor of two. ## Teachers’ and Students’ Experiences - Teachers report higher student engagement and enthusiasm in class. - Students with no previous exposure to design have become interested in pursuing it after graduation. - Some students use Figma at home for personal projects simply because they enjoy creating. - Teachers use FigJam for icebreakers, collaborative exercises, and classroom activities. ## Learning Through Play and Projects - Dylan Field emphasizes play, hands-on work, and project-based learning as important ways to develop design skills. - He describes his own project-based high school, where science concepts were reinforced through engineering projects. - That curriculum reflected a design process: - Explore broadly and generate possibilities - Narrow the options and choose a direction - Execute the selected idea - This approach helps students learn by making, experimenting, and solving concrete problems rather than relying solely on instruction. Figma’s Chromebook initiative is intended to make design education a normal part of students’ everyday learning. Providing free, accessible tools—combined with hands-on, collaborative projects—can help more students build the creative and technical skills needed for the future.

Read original(opens in new tab)
datadog3 min readCurated summary

Engineering Spotlight: Tay Nishimura

Tay Nishimura’s career shows that succeeding in tech is often less about fitting a standard engineering mold and more about finding work that matches one’s strengths. Although she initially struggled with the speed and coding demands of software development, her rigor, visual thinking, and careful approach became valuable in site reliability engineering. Her transition was enabled by self-directed learning, community education, and ToyNet, an open source networking simulator that demonstrated her practical abilities. ## Entering Tech from Mathematics - Tay began as a mathematics major focused on real analysis, then added computer science after advice from a professor. - Internships at Amazon and Google introduced her to the technology industry. - She found a sharp contrast between academia and industry: - School rewarded theoretical rigor. - Industry emphasized practical, fast, and agile solutions. - Tay also felt like an outsider because she had little exposure to computers growing up. ## Struggling with Traditional Software Engineering - Coding did not come naturally to Tay’s visual way of thinking. - She translated code into drawings to understand and modify it, then converted those ideas back into code. - This process produced high-quality, careful work but made her slower than colleagues expected. - A manager suggested product management and site reliability engineering as possible alternatives. - Tay discovered that her deliberate pace was useful for SRE work, particularly when evaluating failure modes and making critical changes. - Because her company offered no path into those roles, she eventually left rather than continue facing increasing stress. ## Discovering Networking and Technical Program Work - Tay’s next role had a software engineer title but involved work closer to product or technical program management. - She learned that job titles and responsibilities vary significantly between companies. - With better work-life balance, she began studying computer networking in her free time. - She created visual diagrams and learning modules to explain switches, routers, and packet flows. - These efforts became Project Reclass, a nonprofit teaching technical skills to incarcerated people and military veterans. - The program used improvised equipment, such as fake routers and switches, to teach concepts in environments where real networking hardware was unavailable. ## Building ToyNet During the Pandemic - After her company laid off its entire office during COVID-19, Tay decided to pursue SRE directly. - When prisons suspended in-person education, Project Reclass adapted by creating a digital networking simulator. - Tay architected ToyNet, an open source platform built with: - React - A Flask backend - Containerized Mininet instances for network emulation - Users can connect simulated routers, switches, and hosts, configure IP addresses, and run commands such as `ping` and `arp`. - ToyNet was designed to work for incarcerated learners with restricted internet access. - Deploying it in the cloud also gave Tay practical experience that helped compensate for limited professional cloud experience. - Companies interested in the project were more likely to advance her through the interview process, eventually leading to Datadog. ## Finding the Right Environment at Datadog - At Datadog, Tay learned Kubernetes, chaos engineering, network traffic control, and Go. - She found that her rigor and visual thinking were assets rather than liabilities. - While learning Datadog’s Chaos Controller codebase, she mapped files and dependencies by drawing boxes and arrows. - Her experience suggests that engineers do not need to learn or reason in a single conventional way; the right environment can turn an apparent weakness into a strength. Tay’s path recommends experimenting broadly, studying independently, and building concrete projects that reveal how you think and solve problems. The most suitable tech role may emerge only after moving between companies and disciplines rather than forcing yourself to succeed in an ill-fitting position.

Read original(opens in new tab)
figma2 min readCurated summary

How (and why) we built branching | Figma Blog

Figma built branching to preserve the benefits of real-time collaboration without letting unfinished or experimental work disrupt approved designs. Branches provide isolated spaces for exploration while keeping the main file a reliable source of truth. The design prioritizes simplicity, familiar multiplayer behavior, and protection against data loss. ## The Need for Branching at Scale - Figma’s real-time multiplayer model usually helps teams work in one shared file. - As organizations grow, however, design workflows become harder to manage: - Unapproved changes can reach production code. - Work may be overwritten. - Teams may struggle to distinguish work in progress from approved designs. - Branches let designers experiment, iterate, contribute to libraries, or preview work without changing the main file. - Changes can be incorporated into the main file only after review and approval. ## Balancing Freedom and Structure - Growing design teams increasingly requested stronger version-control workflows. - Traditional branching comes from software development, where files and changes are stored locally. - Figma had to adapt the concept for cloud-based files, asking: - Whether people should continue editing the main file simultaneously. - How collaboration should work within branches. - How to introduce version control without overwhelming designers with complexity. ## Designing Around Figma’s Multiplayer Model - Figma chose a purpose-built approach rather than copying existing development tools. - Simplicity and consistency became core principles: - The main file remains multiplayer. - A branch behaves like a regular Figma file. - Existing editor and viewer permissions continue to apply. - The system emphasizes data safety so users’ changes remain protected. - Figma intentionally limited complexity, including preventing branches from being created from other branches. ## The Complexity of Merging - Merging involved more than simply resolving conflicting edits. - Figma also had to handle operational problems such as: - The file changing while someone reviews a merge. - A user losing their connection midway through the process. - These cases required safeguards to ensure merges remain understandable, recoverable, and safe. Branching is therefore presented as a structured layer on top of Figma’s collaborative model: teams can explore freely in branches while maintaining a dependable, approved main file.

Read original(opens in new tab)
figma3 min readCurated summary

Why Professors at Stanford and UC Berkeley Use Figma to Teach Design | Figma Blog

Figma argues that interface design courses need tools that minimize technical overhead so professors can focus on design principles. Its browser-based, collaborative platform is presented as accessible across devices, free for educational institutions, and easy for beginners to learn. The post concludes that these features make teaching, feedback, sharing, and group work more efficient. ## The Challenge of Teaching Interface Design - Interface design remains uncommon in college curricula, often appearing as isolated interdisciplinary courses. - Professors must teach substantial technical concepts in limited time. - Traditional design tools can be difficult and time-consuming for students to learn. ## Accessible on Any Computer - Figma runs in the browser and supports PCs, Macs, Linux systems, and Chromebooks. - Students can work anywhere with internet access, without downloading large software packages. - This helps students who do not own high-powered personal computers. ## Free for Educational Institutions - Figma’s premium version is offered free to educational institutions. - Schools do not need to negotiate expensive contracts. - The lack of cost reduces barriers for students from lower-income backgrounds. ## Easy to Learn - Figma is designed around the core functionality digital designers need. - Berkeley instructors describe it as lightweight, simple, and intuitive. - Students with no prior design experience can reportedly learn the basics in under an hour. - Professors and TAs can concentrate on design principles rather than tool training. ## Technical Support for Classes - Figma provides in-app chat support for students and instructors. - The company states that it responds within 24–48 hours on weekdays. - This reduces the technical-support burden on professors and TAs, especially in large classes. ## Always-Current, Shareable Designs - Designs are stored online and accessed through a file URL. - Instructors can view the latest version without requiring students to export or email files. - Teachers can review work or watch students design in real time. ## Feedback and Version History - Comments can be pinned directly to specific design frames. - Students can view feedback immediately and iterate without waiting for in-person meetings. - Version history shows how a design developed and helps instructors assess contributions in group projects. ## Real-Time Collaboration - Multiple students can work in the same file simultaneously. - Cloud collaboration eliminates the need to email files, manage conflicting versions, or take turns editing. - The post also introduces seamless developer handoff as another benefit, though the provided excerpt does not elaborate on it. Figma’s combination of low cost, cross-platform access, simple onboarding, collaboration, and built-in feedback makes it a practical classroom tool for teaching design fundamentals rather than software mechanics.

Read original(opens in new tab)