Cloud Native

8 posts

gitlab2 min readCurated summary

Introducing the 2026 EMEA GitLab Partner Award winners

GitLab announced its 2026 EMEA Partner Award winners, recognizing organizations that drove customer success, technical innovation, certification, business growth, and joint marketing. The awards highlight partners helping enterprises adopt DevSecOps, cloud-native platforms, managed services, and AI-enabled software development across the region. ## Regional Partners of the Year - **Central Europe: cc cloud GmbH** — Combines infrastructure and DevOps expertise to manage cloud applications, platforms, and IT operations. - **Northern Europe: Eficode** — Supports more than 1,600 customers through consulting, managed services, toolchain implementation, and AI-augmented development. - **Southern Europe: Kiratech** — Helps enterprises modernize infrastructure using cloud-native, DevOps, and PlatformOps practices. - **Eastern Europe and Israel: Bynet** — An established systems integrator supporting enterprise IT, cloud, cybersecurity, modernization, DevSecOps, and AI adoption. ## Technical and Enablement Awards - **Best Technical Solution/Project: Capgemini | Sogeti** — Recognized for impactful, complex technical solutions using AI-driven quality engineering, data, and cloud capabilities. - **Most Certified and Enabled Partner: Devoteam** — Awarded for having the largest number of GitLab-certified professionals. - **Rookie of the Year: ITDOTCOM** — A Uzbekistan-based technology distributor that achieved rapid success supporting software, infrastructure, cybersecurity, and business automation across Central Asia. ## Growth and Collaboration Awards - **First Order Master: Linux Polska** — Recognized for winning new customers and business through open-source consulting, DevOps, automation, containerization, and data analytics. - **Co-marketing Partner of the Year: Conoa, a PROACT Company** — Honored for joint marketing efforts and expertise in Kubernetes, cloud-native technologies, container platforms, and managed operations. The awards demonstrate the breadth of GitLab’s EMEA partner ecosystem, from regional systems integrators and cloud specialists to technical consultants and Kubernetes providers. Together, these partners are helping customers modernize delivery practices and adopt DevSecOps and AI capabilities.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: AWS DevOps Agent & Security Agent GA, Product Lifecycle updates, and more (April 6, 2026) | Amazon Web Services

The April 6, 2026 AWS Weekly Roundup highlights the general availability of AWS DevOps Agent and AWS Security Agent, autonomous “frontier agents” designed to handle complex operational and security tasks. It also reviews AWS service lifecycle changes and summarizes notable product launches and technical updates from the previous week. The overall message is that AWS is expanding agentic automation while helping customers manage service transitions and adopt new capabilities. ## AWS DevOps Agent and Security Agent Reach GA - **AWS DevOps Agent** - Investigates incidents, accelerates resolution, and helps prevent recurring problems. - Works continuously across multiple steps until an operational goal is complete. - Customers report up to **75% lower mean time to resolution (MTTR)** and **3–5 times faster incident resolution**. - Western Governors University reduced resolution times from hours to minutes. - **AWS Security Agent** - Provides continuous, context-aware penetration testing during the software development lifecycle. - Operates similarly to a human penetration tester. - LG CNS reported testing that was more than **50% faster**, approximately **30% less expensive**, and produced fewer false positives. - **Deployment flexibility** - Both agents support AWS, multicloud, and on-premises environments. - They are intended to automate repetitive investigative and testing work while allowing teams to focus on higher-value activities. ## AWS Service Lifecycle Changes AWS updated its Product Lifecycle Changes guidance on March 31, 2026, including migration recommendations and alternative services. - Services with availability changes or maintenance guidance include: - AWS App Runner - AWS Audit Manager - AWS CloudTrail Lake - AWS Glue Ray jobs - AWS IoT FleetWise - Amazon Application Recovery Controller Readiness Check - Amazon Comprehend features such as Topic Modeling and Prompt Safety Classification - Amazon Rekognition streaming and batch moderation features - Amazon SNS Message Data Protection - Services listed as entering sunset include: - AWS Service Management Connector - Amazon RDS Custom for Oracle - Amazon WorkMail - Amazon WorkSpaces Thin Client - **Amazon Chime SDK Proxy Sessions** is reaching sunset. AWS recommends reviewing the relevant service documentation or contacting Support to reduce operational disruption. ## Notable AWS Launches - Amazon ECS introduced **Managed Daemons for ECS Managed Instances**. - The AWS Sustainability console now consolidates **Scope 1–3 emissions reporting**. - **Amazon Bedrock AgentCore Evaluations** became generally available. - AWS Transform added generally available automated codebase analysis. - CloudWatch introduced OpenTelemetry Container Insights for Amazon EKS in preview. - Amazon Lightsail added compute-optimized bundles with up to **72 vCPUs**. - Amazon CloudFront added **SHA-256 support** for signed URLs and signed cookies. ## Additional AWS Resources The roundup also points readers to material on: - Architecting agentic AI applications on AWS. - Reducing data-transfer costs with Network Load Balancers. - Preventing hallucinations in production AI agents. - The AWS World Sports Innovation Cup. - Exploring AWS communities through an interactive 3D globe. AWS also encourages readers to participate in Builder Center discussions, community events, AWS Summits, and developer-focused programs. AWS teams should review the lifecycle notices for services they depend on, while developers and operations groups may benefit from evaluating the new agents and launches for automation, security testing, and observability improvements.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: Amazon Connect Health, Bedrock AgentCore Policy, GameDay Europe, and more (March 9, 2026) | Amazon Web Services

The March 9, 2026 AWS Weekly Roundup highlights AWS’s growing focus on agentic AI, healthcare automation, security, and developer productivity. Major updates include Amazon Connect Health, centralized policies for Bedrock agents, private AI assistants on Lightsail, and new tools for troubleshooting and durable Lambda workflows. The roundup also previews community events, including GameDay Europe, NVIDIA GTC, AWS Summits, and regional Community Days. ## Major AWS Product Launches - **Amazon Connect Health** is generally available with five healthcare-focused AI agents: - Patient verification - Appointment management - Patient insights - Ambient documentation - Medical coding - These capabilities are HIPAA-eligible and designed to integrate with existing clinical workflows within days. - **Bedrock AgentCore Policy** provides centralized, fine-grained controls for agent-to-tool interactions. - Policies can be written in natural language. - AWS converts them into Cedar, its open-source policy language. - Controls operate outside application code, supporting security and compliance teams. - **OpenClaw on Amazon Lightsail** enables deployment of private autonomous AI assistants. - Includes sandboxed sessions, security controls, HTTPS, and device-pairing authentication. - Uses Amazon Bedrock by default and supports Slack, Telegram, WhatsApp, and Discord integrations. ## Pricing, Cost Management, and Security - **VPC Encryption Controls** became a paid feature on March 1, 2026. - Monitor mode detects unencrypted traffic. - Enforce mode blocks traffic that does not meet encryption requirements. - Controls apply to traffic within and across VPCs in a region. - **Database Savings Plans** now cover Amazon OpenSearch Service and Amazon Neptune Analytics. - Customers can save up to 35% with a one-year commitment. - Savings apply across engine, instance family, size, and AWS Region. - **Amazon GameLift Servers DDoS Protection** adds a co-located relay network. - Client traffic is authenticated with access tokens. - Per-player traffic limits help mitigate attacks. - The feature adds no cost for GameLift Servers customers. ## Developer and Operations Improvements - **Elastic Beanstalk AI-powered environment analysis** sends events, health data, and logs to Amazon Bedrock when environments degrade. - It returns troubleshooting recommendations tailored to the affected environment. - AWS now allows **IAM roles to be created directly inside service workflows**, reducing the need to switch to the IAM console. Supported services include EC2, Lambda, EKS, ECS, Glue, and CloudFormation. - **Kiro’s new Lambda durable functions power** assists developers with long-running, multi-step applications and AI workflows. - It provides guidance on replay models, waits, concurrency, error handling, and deployment. ## AWS Community Projects - One community project demonstrates a persistent AI memory layer using **MCP, Amazon Bedrock, and a Chrome extension**, allowing agents to retain context across sessions and applications. - Another experimental application treats the AI model as the runtime, generating a complete interactive web application from a single prompt without a conventional codebase, framework, or persistent state. ## Community Events and AWS Activities - **AWS Community GameDay Europe** takes place March 17, offering team-based challenges using real AWS services. - AWS will participate in **NVIDIA GTC 2026** in San Jose from March 16–19, with sessions, demos, booths, and discounted passes. - Upcoming **AWS Summits** include Paris, London, and Bengaluru. - Upcoming **AWS Community Days** include events in Slovakia, Pune, and Mexico City. AWS’s latest announcements show a clear emphasis on practical AI agents, stronger governance, and automation across infrastructure and application development. Developers and cloud teams should review the new security and pricing changes while exploring the AI tools and upcoming hands-on community events.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Complexity is a choice. SASE migrations shouldn’t take years.

Cloudflare argues that SASE and zero trust migrations do not need to take years. Its partners, TachTech and Adapture, reportedly reduced deployments from around 18 months to four–six weeks by using Cloudflare One’s unified, cloud-native architecture. The post concludes that programmable security infrastructure can accelerate zero trust adoption while also enabling safer use of AI. ## Faster Zero Trust Deployments - Traditional Secure Web Gateway (SWG) and Zero Trust Network Access (ZTNA) migrations can take up to 18 months for large organizations. - TachTech reduced comparable Cloudflare One deployments to four–six weeks. - Cloudflare Access is presented as lightweight and largely “no-touch” after deployment, reducing ongoing operational effort. ## Why Legacy Migrations Stall - Legacy architectures often treat migration as hardware replacement rather than software transformation. - Complex service chaining creates a “trombone effect,” increasing latency and making troubleshooting difficult. - Cloudflare’s partners accelerate migrations through: - **Identity-first on-ramps:** Existing identity-provider groups define access policies instead of rebuilding network segments. - **Consolidated policy engines:** SWG and ZTNA policies are handled together, avoiding synchronization between separate products. - **Cloud-native connectors:** Tools such as `cloudflared` provide connectivity without opening inbound firewall ports. ## Scaling Quickly - Adapture expanded one Cloudflare Access deployment from 600 contractors to 5,000 users. - The company describes the expansion as seamless compared with the lengthy implementation cycles associated with legacy SASE platforms. - Cloudflare positions rapid elasticity as important for organizations whose workforce and security needs change quickly. ## A Programmable, Extensible Edge - Cloudflare One is described as software-defined and composable, allowing partners to adapt it to specialized environments. - TachTech supported Arch Linux developer workstations by extracting binaries from an Ubuntu `.deb` package and creating a custom `PKGBUILD`. - This approach preserved device-posture checks, including disk-encryption and firewall-status verification, without creating a security exception. ## Supporting Safe AI Adoption - Cloudflare says the Secure Web Gateway is evolving from simple URL filtering toward controlling data flows to large language models. - Its AI security capabilities include: - **Shadow AI visibility:** Identifying unauthorized AI tools in use across the organization. - **AI confidence scores:** Evaluating models based on standards such as SOC 2 and ISO 42001, as well as data-handling practices. - **DLP prompt protection:** Blocking sensitive source code, personally identifiable information, and financial data from being submitted to public AI services. - **LLM discovery:** Finding and labeling internet-exposed LLM endpoints to reveal the organization’s AI attack surface. - **Request validation:** Intended to defend AI applications against prompt injection and related attacks. Cloudflare’s central recommendation is to replace fragmented, hardware-oriented security deployments with a unified, programmable platform. Doing so can shorten zero trust migrations, simplify operations, preserve consistent security controls across unusual environments, and establish a faster foundation for responsible AI adoption.

Read original(opens in new tab)
gitlab4 min readCurated summary

How GitLab built a security control framework from scratch

GitLab created its own security control framework after finding that existing frameworks were too broad, rigid, or insufficiently granular for its multi-product, cloud-native environment. The GitLab Control Framework (GCF) combines industry best practices with product-specific implementations and extensive operational metadata. This lets GitLab manage multiple certifications and internal risks through one scalable framework rather than maintaining separate frameworks for each product. ## Why Existing Frameworks Were Insufficient - GitLab initially used the Secure Controls Framework, then adopted NIST SP 800-53 in preparation for FedRAMP. - NIST’s more than 1,000 controls were comprehensive but included requirements that did not apply to GitLab. - Broad controls often combined several distinct activities: - NIST AC-2, “Account Management,” covers account creation, modification, disabling, termination, shared accounts, and monitoring. - GitLab treated these as separate controls because they have different owners, risks, testing methods, and evidence requirements. - Repeatedly customizing NIST controls effectively meant GitLab was building its own framework, leading to the decision to formalize one. ## Establishing the GitLab Control Framework GitLab developed the GCF through five major steps: ### Assessing Requirements - The team mapped requirements from existing and planned certifications, including: - SOC 2 Type II - ISO 27001, ISO 27017, ISO 27018, and ISO 42001 - PCI DSS - TISAX - Cyber Essentials - FedRAMP - Internal requirements covered mission-critical systems outside certification scopes and systems handling sensitive data. - This analysis established the minimum controls GitLab needed to meet compliance and risk-management obligations. ### Learning from Industry Frameworks - GitLab compared its requirements with: - NIST SP 800-53 - NIST Cybersecurity Framework - Secure Controls Framework - Adobe and Cisco Common Controls Framework - The goal was to reuse proven structures and ensure important security domains and practices were not omitted. ### Creating Custom Domains - The team organized the framework into 18 custom control domains. - Each domain groups related controls according to how GitLab’s security program is managed. - The structure supports adding, changing, or retiring controls as the business evolves. ## Separating Framework Requirements from Implementations GitLab operates several products with different infrastructure and compliance scopes: - GitLab.com is a multi-tenant SaaS platform hosted on GCP. - GitLab Dedicated is single-tenant SaaS hosted on AWS. - GitLab Dedicated for Government is a FedRAMP offering hosted on AWS. To avoid duplicating the framework, the GCF uses two control levels: - **Level 1:** Defines what must be implemented at the organizational framework level. - **Level 2:** Describes how each product fulfills the requirement. - Entity-level controls apply across the organization and are inherited by all product offerings. - This model supports product-specific audits while preserving a single source of control requirements. ## Adding Operational Metadata Rather than tracking only a control ID, description, and owner, the GCF records detailed context for each control: - Responsible owner and risk accountability - Applicable environment or product - Covered assets and systems - Performance or testing frequency - Manual, semi-automated, or automated nature - External certification or internal-risk classification - Testing procedures and required evidence This turns the framework into an operational control inventory. Teams can filter it to identify controls for a particular audit, determine ownership, or find manual controls that may be candidates for automation. ## Designing for Growth - The GCF is intended to evolve with GitLab’s products, risks, and certification goals. - Its structured metadata helps GitLab assess scope and identify gaps when pursuing additional certifications such as ISMAP, IRAP, or C5. - The framework’s modular design makes it easier to extend compliance coverage without creating entirely new control systems. GitLab’s experience suggests that organizations should consider a custom framework when standard frameworks require extensive modification. The most effective approach is to retain useful industry guidance while tailoring control granularity, product implementations, ownership, testing, and metadata to the organization’s actual operating environment.

Read original(opens in new tab)
netflix3 min readCurated summary

Automating RDS Postgres to Aurora Postgres Migration

Netflix standardized on Amazon Aurora PostgreSQL after finding that PostgreSQL already supported most relational workloads and that Aurora offered stronger scalability, availability, elasticity, and ecosystem alignment. To migrate nearly 400 RDS PostgreSQL clusters efficiently, Netflix built a self-service workflow that automates replication, traffic quiescence, validation, and cutover while minimizing downtime and eliminating data loss. The Aurora read-replica method is preferred over snapshot migration because it keeps the target nearly synchronized while production continues running. ## Why Netflix Chose Aurora PostgreSQL - PostgreSQL already supported the majority of Netflix’s relational workloads. - Internal evaluations found Aurora PostgreSQL could support more than 95% of workloads running on other relational database systems. - PostgreSQL benefits from: - A broad open-source ecosystem - Strong community adoption - Compatibility with modern data platforms - Aurora’s distributed, cloud-native architecture provides: - Better scalability and elasticity - High availability - Support for globally distributed applications - The migration effort began with RDS PostgreSQL and is intended to expand to other relational systems. ## Database Migration Requires More Than Data Copying A safe migration must move both data and database functionality while preserving correctness, availability, and performance. - **Data replication:** Copy existing data and continuously apply source changes to the destination. - **Quiescence:** Stop writes to the source so the destination can catch up completely. - **Validation:** Confirm that source and destination data are synchronized. - **Cutover:** Redirect applications to the new Aurora database as the system of record. ## Operational and Technical Challenges - Manually migrating almost 400 PostgreSQL clusters would be slow, error-prone, and operationally expensive. - Coordinating downtime across dependent services is difficult. - Netflix therefore created a self-service workflow that handles orchestration, safety checks, and correctness guarantees automatically. - The system must guarantee: - Zero data loss - Extremely short downtime, especially for critical services - No performance degradation during or after migration - Migration of related resources such as parameter groups, read replicas, and replication slots - Application teams control database clients, so the platform cannot depend on them manually pausing writes. - The migration system must provide control-plane mechanisms to halt traffic safely during validation and cutover. - The workflow must operate without obtaining RDS credentials from users, since databases may be tightly secured and the migration platform may lack direct database access. - Because non-experts operate the process, the experience must be self-guided and require minimal user effort. ## Snapshot-Based Migration The snapshot approach is straightforward but requires stopping writes before migration. - Halt write traffic to the RDS PostgreSQL source. - Create a manual snapshot. - Convert the snapshot into an Aurora-compatible format. - Create an Aurora PostgreSQL cluster from the converted snapshot. - Validate the new cluster. - Redirect applications to the Aurora endpoint. This method is simple but can involve a longer interruption because the target is not continuously updated while the snapshot is created and converted. ## Aurora Read-Replica Migration The read-replica approach reduces downtime by continuously replicating the RDS database into Aurora. - Create an Aurora PostgreSQL read replica from the RDS source. - Stream changes asynchronously from RDS to Aurora while applications continue using the source. - Provision and validate Aurora configuration, connectivity, and performance in advance. - When replication lag is sufficiently low, briefly pause writes. - Allow the replica to catch up fully. - Promote it to a standalone Aurora PostgreSQL cluster. - Redirect application traffic to the Aurora endpoint. This approach keeps the destination nearly synchronized before cutover, making it substantially less disruptive than snapshot-based migration. Netflix’s automation focuses on making the read-replica migration process safe, repeatable, and self-service, with the platform handling replication, traffic control, validation, and cutover rather than relying on manual application-team coordination.

Read original(opens in new tab)
lineOriginal article

Why an Athenz Engineer Took (opens in new tab)

Security platform engineer Jung-woo Kim details his transition from a specialized Athenz developer to a "Kubestronaut," a prestigious CNCF designation awarded to those who master the entire Kubernetes ecosystem. By systematically obtaining five distinct certifications, he argues that deep, practical knowledge of container orchestration is essential for building secure, scalable access control systems in private cloud environments. His journey demonstrates that moving beyond application-level expertise to master cluster administration and security directly improves architectural design and operational troubleshooting. ## The Kubestronaut Framework * The title is awarded by the Cloud Native Computing Foundation (CNCF) to individuals who pass five specific certification exams: CKA, CKAD, CKS, KCNA, and KCSA. * The CKA (Administrator), CKAD (Application Developer), and CKS (Security Specialist) exams are performance-based, requiring candidates to solve real-world technical problems in a live terminal environment rather than answering multiple-choice questions. * Success in these exams demands a combination of deep technical knowledge, speed, and accuracy, as practitioners must configure clusters and resolve failures under strict time constraints. * The remaining Associate-level exams (KCNA and KCSA) provide a theoretical foundation in cloud-native security and ecosystem standards. ## A Progressive Path to Technical Mastery * **CKAD (Application Developer):** The initial focus was on mastering the deployment of Athenz—an open-source auth system—ensuring it runs efficiently from a developer's perspective. Preparation involved rigorous use of tools like killer.sh to simulate high-pressure environments. * **CKA (Administrator):** To manage multi-cluster environments and understand the underlying components that make Kubernetes function, the author moved to the administrator level, gaining insight into how various services interact within the cluster. * **CKS (Security Specialist):** Given his background in security, this was the most critical and difficult stage, focusing on cluster hardening, vulnerability analysis, and implementing strict network policies to ensure the entire infrastructure remains resilient. ## Organizational Impact and Open Source Governance * Obtaining these certifications provided a clearer understanding of open-source governance, specifically how Special Interest Groups (SIGs) and pull request (PR) workflows drive massive projects like Kubernetes. * This technical depth was applied to a high-stakes project providing Athenz services in a Bare Metal as a Service (BMaaS) environment, allowing for more stable and efficient architecture design. * The learning process was supported by corporate initiatives, including access to Udemy Business for technical training and a hybrid work culture that allowed for consistent, early-morning study habits. To achieve expert-level proficiency in complex systems like Kubernetes, engineers should adopt the "Ubo-cheonri" philosophy—making slow but steady progress. Starting with even one minute of study or a single GitHub commit per day can eventually lead to mastering the highest levels of cloud-native architecture. For those managing enterprise-grade infrastructure, pursuing the Kubestronaut path is highly recommended as it transforms theoretical knowledge into a broad, practical vision for system design.

figma2 min readCurated summary

Our approach to security at speed | Figma Blog

Figma’s security team aims to help teams ship quickly without compromising safety. Its approach combines early risk assessment, reusable technical controls, decentralized decision-making, and transparent collaboration rather than rigid mandates. The goal is to make security an enabler of product development. ## Systematically Assessing Risk - Security reviews upcoming features and workflows to identify risks early. - Teams use a three-question “ThreatJam” survey before security office hours, held three times weekly. - For FigJam’s rich link previews, the team identified risks including: - **SSRF**, where attackers could request internal Figma resources. - **Denial-of-service attacks** against linked websites. - Malicious HTML that could deface or execute code within FigJam. - Figma isolated link scraping and parsing in a cloud function running on a separate virtual machine and network. - Additional protections included: - Cloud-function rate limiting. - Temporary storage of scraped data. - Restricting requests to standard HTTP and HTTPS ports. - Limiting parsed HTML to approved tags. - Reusing existing plugin and widget security mechanisms. ## Reusable and Decentralized Security Solutions - Security provides customized guidance while avoiding centralized approval processes. - The team builds reusable libraries, frameworks, documentation, and security patterns. - Solutions developed for one product can serve as case studies for other engineering teams. - Office-hours notes are shared across Figma so teams can understand security reasoning and apply it independently. - This model allows security practices to scale as the company grows. ## Defending Against Phishing and Information Disclosure - Figma uses technical controls to protect employees, devices, and internal data. - Access to internal sites requires a passwordless second factor scoped to the specific site. - New employees receive hardware authenticator keys. - Employees are also encouraged to register biometric authenticators such as Touch ID or Windows Hello. Figma’s model demonstrates that security can move at development speed when teams assess threats early, isolate risky functionality, automate defenses, and share solutions broadly. Companies seeking a similar approach should prioritize reusable controls and security collaboration over process-heavy gates.

Read original(opens in new tab)