Google Cloud

9 posts

gitlab2 min readCurated summary

GitLab on Google Cloud: Fully managed, compliant, and AI-ready

GitLab is introducing a fully managed deployment on Google Cloud through certified managed service providers such as Beyond and Digital Future. The offering combines data residency and compliance controls with access to Google’s Gemini and Gemma models through GitLab Duo Agent Platform. Organizations can also purchase the platform through Google Cloud Marketplace, applying existing cloud commitments to GitLab, AI inference, and infrastructure costs. ## Fully Managed GitLab on Google Cloud - Certified MSPs operate GitLab on Google Cloud under service-level agreements, removing infrastructure-management responsibilities from customer teams. - Organizations retain control over where code, pipelines, and security data are stored, supporting sovereignty and data-residency requirements. - GitLab’s audit and policy controls provide visibility into agent actions, merge requests, and security findings. ## AI Model Choice for Different Workloads - Gemini models, including Gemini 3.5 Flash, are available in Duo Agent Platform through Google’s Gemini Enterprise Agent Platform. - GitLab’s participation in Google’s early-access program is intended to bring new Gemini models to Duo as they become available. - Regulated or self-hosted teams can use Gemma 4 with GitLab Duo Self-Hosted. - With self-hosted models, the AI Gateway and all requests and responses remain within an organization’s on-premises or private-cloud environment. ## Using Existing Google Cloud Commitments - GitLab and Duo Agent Platform can be purchased through Google Cloud Marketplace. - Existing Google Cloud commitments can fund GitLab subscriptions, model inference, and related infrastructure without starting a new procurement cycle. - Consolidated Google Cloud billing reduces reconciliation across vendors. - GitLab retains its own cost-management features, including usage dashboards, model policies, and GitLab Credits for more predictable AI spending. ## One Governed DevSecOps Platform - GitLab Duo Agent Platform provides software-delivery context that standalone coding assistants lack, including merge requests, pipelines, and deployment targets. - This context helps agents perform multi-step work and supports code review at monorepo scale. - Combining GitLab’s governance and lifecycle data with Google’s models keeps deployment, model selection, compliance, and spending aligned in one platform rather than fragmented across multiple tools. Organizations can start with a Duo Agent Platform trial, enable it through the free GitLab tier, or use included GitLab Credits with Premium and Ultimate subscriptions. Overall, the offering is aimed at teams that want managed GitLab operations, flexible AI model access, and strong control over data location and costs on Google Cloud.

Read original(opens in new tab)
gitlab3 min readCurated summary

GitLab AI Hackathon 2026: Meet the winners

Nearly 7,000 developers participated in GitLab’s 2026 AI Hackathon, creating more than 600 agents and workflows for the GitLab Duo Agent Platform. The winning projects focused on practical software delivery challenges—including organizational knowledge loss, security, compliance, migrations, and sustainability—rather than simple chatbot interactions. The results suggest that agentic AI is becoming most valuable when integrated directly into development workflows and given richer project context. ## Hackathon Scope and Evaluation - The hackathon ran from February 9 to March 25, 2026, on Devpost. - Google Cloud and Anthropic co-sponsored the event, contributing judges, prizes, and cloud resources. - Nineteen judges evaluated projects on: - Technical execution - Design - Potential impact - Quality of the idea - Total prizes reached $65,000. ## Grand Prize: LORE - LORE, or Living Organizational Record Engine, addresses the loss of institutional knowledge when engineers leave. - It combines: - Eight specialized agents - A router that directs questions to the appropriate agent - Protections against circular loops in its knowledge graph - A visual dashboard - Carbon tracking - Its command-line tool includes 43 tests, leading judges to describe it as a polished product rather than a typical hackathon prototype. ## Google Cloud and Anthropic Winners - **Gitdefender**, the Google Cloud Grand Prize winner, detects security issues during code review, writes fixes, and opens the review automatically. - **Aegis**, the Google Cloud Runner Up, explains the reasoning behind its AI decisions and is deployed on Google Cloud. - **GraphDev**, the Anthropic Grand Prize winner, maps code relationships and shows how systems evolve, helping developers understand the impact of changes. - **DocSync**, the Anthropic Runner Up, uses Detector, Writer, and Reviewer agents to update documentation. It opens a review when confident and creates an issue for human review when uncertain. ## Category Winners - **Time-Traveler**, winner for technical achievement, creates a safe copy of a production environment and runs database migrations against it using five connected agents, PostgreSQL, real data, and Google Cloud deployment. - **RedAgent**, the most impactful project, verifies AI-generated security findings before developers act on them, addressing distrust in automated reports. - **Launch Control**, recognized for ease of use, combined polished user experience, strong infrastructure, and sustainability considerations. ## Sustainability-Focused Projects - Five projects received sustainability prizes or bonuses as the organizers highlighted the growing energy cost of CI/CD systems and large language models. - **GreenPipe** analyzes CI/CD pipelines and generates carbon-footprint reports. - Sustainable Design bonuses recognized projects including: - **BugFlow**, which generated 10 fixes from one bug report in 20 minutes - **DELTA Cyber Reasoning**, an automated fuzz-testing tool - **CarbonLint**, which applies code analysis to energy consumption - **TFGuardian**, which includes carbon-footprint analysis - One project reduced monthly costs from $556 to $18, representing a reported 96% carbon reduction. ## Honorable Mentions - **SecurityMonkey** tests security scanners by injecting known vulnerabilities. - **stregent** enables CI/CD investigation and fixes through WhatsApp. - **Compliance Sentinel** evaluates merge requests for compliance risk and blocks critical violations. - **Carbon Tracker** measures the carbon footprint of individual pipeline jobs and suggests improvements. - **RepoWarden** captures the rationale behind code, not only its behavior. - **MR Compliance Auditor** maps merge-request evidence to SOC 2 controls and displays compliance scores in real time. ## What Comes Next The projects operated within a single GitLab project, but many teams supplemented their agents with local knowledge graphs to understand code relationships and dependencies. GitLab plans to build on this approach in future hackathons by providing agents with richer context. GitLab’s hackathon demonstrates that the strongest AI agents are workflow-integrated tools that can investigate, make decisions, execute changes, and involve humans when needed. Developers can explore the 600-plus projects in the gallery or build their own agents on the GitLab Duo Agent Platform.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Interconnect is now generally available, with a new option to simplify last-mile connectivity | Amazon Web Services

AWS Interconnect is a managed service for private, high-speed connectivity between AWS and other clouds or on-premises networks. Its multicloud capability, now generally available, initially connects AWS with Google Cloud while Microsoft Azure support is planned for later in 2026. The service aims to replace complex VPN, colocation, and third-party networking setups with a turnkey, resilient configuration managed through AWS. ## AWS Interconnect Capabilities - **Interconnect – multicloud** connects an AWS VPC privately to VPCs on other cloud providers. - **Interconnect – last mile** simplifies connectivity from branch offices, data centers, and remote sites through existing network providers. - Both capabilities provide: - Dedicated bandwidth - Private connectivity - Managed provisioning - Reduced infrastructure and configuration overhead - Connections can be configured through the AWS Console by selecting the location or provider, AWS Region, and bandwidth. ## Multicloud Connectivity - The service provides a managed **Layer 3 connection** between AWS and another cloud provider. - Traffic uses the AWS global backbone and the partner’s private network rather than the public internet. - This improves: - Latency predictability - Throughput consistency - Isolation from internet congestion - Google Cloud is supported at launch; Microsoft Azure is expected later in 2026. ## Security, Resilience, and Monitoring - Physical links between AWS and partner routers use **IEEE 802.1AE MACsec encryption** by default. - Each cloud provider handles encryption on its own backbone, so customers must verify that the resulting deployment satisfies compliance requirements. - Connections use multiple logical links across at least two physical facilities to protect against device or facility failures. - Amazon CloudWatch integration includes: - A Network Synthetic Monitor for round-trip latency and packet loss - Bandwidth utilization metrics for capacity planning ## Open Partner Specification - AWS has published the underlying Interconnect specification on GitHub under the **Apache 2.0 license**. - Other cloud providers can become partners by implementing the specification and meeting AWS requirements for: - Resiliency - Support - Service-level agreements - Operational readiness ## Provisioning an AWS–Google Cloud Connection - The demonstration connects a single AWS VPC to a Google Cloud VPC using a Direct Connect Gateway. - In the AWS Direct Connect console, the user: - Selects Google Cloud as the provider - Chooses AWS Region `eu-central-1` - Chooses Google Cloud Region `europe-west3` - Specifies bandwidth - Selects a Direct Connect Gateway - Enters the Google Cloud project ID - AWS then generates an activation key for use on the Google Cloud side. ## Configuring Google Cloud - Because a Google Cloud web console option was unavailable at the time, the example uses the `gcloud` CLI. - The user creates a transport resource with: - The AWS activation key - The Google Cloud region - The target VPC network - Advertised AWS routes - After the transport reaches the appropriate state, the user creates a VPC peering connection between the Google Cloud VPC and the generated transport network. - Custom routes are imported and exported through the peering configuration. ## Completing the AWS Configuration - Once the Google Cloud transport and peering are configured: - The AWS Interconnect status can be checked in the Interconnect console. - The Direct Connect Gateway shows the new attachment. - The final AWS-side step is associating the gateway with the appropriate Virtual Private Gateway. - The Virtual Private Gateway must be in the same AWS Region as the Interconnect. - AWS routing still requires a final route entry so workloads can reach the remote Google Cloud network. AWS Interconnect is best suited to organizations operating hybrid or multicloud environments that want private, resilient connectivity without managing physical links or complex third-party networking. The managed provisioning process can reduce setup time to minutes, but teams should still validate routing, encryption responsibilities, regional constraints, and compliance requirements.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab and Vertex AI on Google Cloud: Advancing agentic development

GitLab is partnering with Google Cloud to combine the GitLab Duo Agent Platform’s lifecycle-wide orchestration with Vertex AI’s managed foundation models and enterprise controls. The integration gives development teams context-aware agents for planning, coding, security, and delivery while keeping workflows within GitLab’s governed system of record. Customers gain model flexibility, stronger governance, and reduced complexity compared with managing disconnected AI tools. ## Agents Across the Software Development Lifecycle - GitLab Duo Agent Platform coordinates specialized agents across planning, development, code review, security, and delivery. - Unlike standalone coding assistants, GitLab agents can access issues, merge requests, pipelines, vulnerabilities, and codebases. - GitLab Duo Planner Agent can analyze backlogs, divide epics into tasks, and support prioritization. - Security Analyst Agent can triage vulnerabilities, explain risks, and recommend remediation priorities. - Built-in flows connect agents into end-to-end processes, reducing manual handoffs. - Agentic Chat provides natural-language access to project context and multi-step reasoning within GitLab. ## Vertex AI as the Model and Infrastructure Layer - Vertex AI supplies the foundation models and related services used by GitLab agents. - Newer models improve reasoning, tool use, and long-context understanding, supporting workloads such as backlog analysis and monorepo security reviews. - Vertex AI Model Garden offers Gemini, third-party, and open-source models, allowing customers to balance performance, cost, and regulatory requirements. - GitLab supports Bring Your Own Model configurations, enabling organizations to use approved providers and gateways. - Vertex AI abstracts LLM hosting, including infrastructure management, security, governance, and model-version delivery. ## Enterprise Governance and Operational Benefits - GitLab’s AI Gateway mediates model access, helping administrators track connections and maintain governance. - Developers remain in GitLab while inference follows existing Google Cloud security and policy controls. - Platform teams can standardize which models support recommendations, analysis, and remediation. - Security teams can manage findings and proposed fixes in the same environment, reducing context switching and unmanaged workflows. - Using Vertex AI through GitLab can align AI usage with existing Google Cloud contracts, controls, and procurement policies. - The approach helps reduce duplicate spending and fragmented “shadow AI” toolchains. ## Practical Outcome for Google Cloud Customers The integration is intended to increase developer productivity without requiring teams to evaluate, host, or manage individual language models. GitLab provides the governed DevSecOps control plane, while Vertex AI supplies scalable, flexible model infrastructure, enabling organizations to adopt more capable agentic workflows while maintaining enterprise security and control.

Read original(opens in new tab)
google3 min readCurated summary

Where wild things roam: Identifying wildlife with SpeciesNet

SpeciesNet is an open-source AI tool that identifies wildlife in camera-trap images, making large-scale monitoring faster and more practical. Trained on more than 65 million labeled images, it can classify nearly 2,500 animal categories and process tens of thousands of images per day. Its adoption by researchers, governments, and conservation groups is expanding wildlife research and enabling more responsive conservation efforts. ## A New Era for Wildlife Monitoring - Motion-triggered camera traps generate enormous volumes of images, often far beyond what human teams can classify manually. - Automated identification helps researchers: - Track population health and changes. - Study migration and climate-related movement. - Estimate population sizes. - Detect rare or endangered species. - SpeciesNet uses deep learning to identify animals in camera-trap photos, accelerating analysis and improving wildlife-management decisions. - The tool is part of Google Earth AI, a collection of geospatial AI tools intended to support environmental and conservation work. ## SpeciesNet’s Training and Performance - SpeciesNet classifies 2,498 categories of mammals, birds, reptiles, and other animals. - It works with MegaDetector, another open-source model that identifies which images and pixels contain animals. - The system provides: - Species names. - Confidence scores. - Multiple identifications when several animals appear in one image. - Processing capacity is approximately: - 30,000 images per day on a standard laptop. - 250,000 or more images per day on a low-end gaming GPU. - SpeciesNet was trained on more than 65 million images from Wildlife Insights and public repositories. - On held-out camera-trap projects, it: - Detected animals in 99.4% of relevant images. - Reached species-level classification 83% of the time. - Produced correct species-level predictions in 94.5% of those cases. - Human-verified labels from Wildlife Insights can be reused as additional training data, creating a feedback loop for improving the model. ## Conservation Projects Using SpeciesNet - **Snapshot Serengeti:** Researchers can analyze roughly 11 million images collected since 2010 in just days, rather than relying exclusively on citizen scientists. Field processing also allows cameras to be redeployed based on recent sightings. - **Wildlife Observatory of Australia:** The organization trained a regional version of SpeciesNet to recognize Australian species missing from the original label set, including musky rat-kangaroos and orange-footed scrubfowl. - **Idaho Department of Fish and Game:** SpeciesNet serves as a first-pass classifier for images of deer, elk, black bears, coyotes, and other wildlife, speeding up human verification. - **Public and private platforms:** Tools including Animl and AddaxAI have integrated SpeciesNet, while companies such as Okala use it alongside Google’s Perch audio model to monitor biodiversity in Africa. - The model has also supported studies of pumas and ocelots in Colombia, cassowaries in Australia, and lions and elephants in Tanzania. SpeciesNet demonstrates how open-source AI can turn massive camera-trap datasets into usable scientific evidence. Its strongest role is as a scalable first-pass system combined with human review, while regional adaptations can extend its usefulness to local and threatened species.

Read original(opens in new tab)
gitlab2 min readCurated summary

Secure and fast deployments to Google Agent Engine with GitLab

Google Agent Engine provides a managed, scalable runtime for AI agents built with Google’s Agent Development Kit (ADK). The post shows how to deploy an ADK agent through GitLab using Workload Identity Federation, avoiding service-account keys while integrating security scanning into CI/CD. A GitLab pipeline can automatically test and deploy the agent to Agent Engine when changes reach the main branch. ## Agent Engine and GitLab - Agent Engine manages infrastructure, scaling, sessions, memory storage, logging, monitoring, and IAM. - GitLab simplifies deployment through: - Dependency scanning, SAST, and secret detection. - Native Google Cloud integration. - Keyless authentication with Workload Identity Federation. - CI/CD templates and the ADK deployment CLI. ## Prerequisites - A Google Cloud project with the Cloud Storage and Vertex AI APIs enabled. - A GitLab project containing the agent source code. - A Google Cloud Storage bucket for deployment staging. - GitLab’s Google Cloud IAM integration configured. ## Configure IAM with Workload Identity Federation - In GitLab, configure the Google Cloud IAM integration with: - Project ID - Project number - Workload Identity Pool ID - Provider ID - Run GitLab’s generated setup script in Google Cloud Shell. - Grant the federated service principal: - `roles/aiplatform.user` - `roles/storage.objectAdmin` - This setup lets GitLab authenticate to Google Cloud without storing long-lived service-account keys. ## Build the GitLab CI/CD Pipeline - Add a `.gitlab-ci.yml` file with `test` and `deploy` stages. - Use the `google/cloud-sdk:slim` image and define variables for: - Google Cloud project and region - Staging bucket - Agent name - Agent entry point - Include GitLab templates for: - Dependency scanning - Static application security testing - Secret detection - Enable keyless authentication with: ```yaml identity: google_cloud ``` - Install the ADK and required Google Cloud libraries during the job. - Deploy with: ```bash adk deploy agent_engine \ --project=$GCP_PROJECT_ID \ --region=$GCP_REGION \ --staging_bucket=gs://$STORAGE_BUCKET \ --display_name="$AGENT_NAME" \ $AGENT_ENTRY ``` - Restrict deployment to the `main` branch. - Cache Python dependencies to speed up later pipeline runs. ## Deploy and Verify - Commit the agent code and `.gitlab-ci.yml` to GitLab. - Monitor the pipeline under **Build > Pipelines**. - Confirm that security scans complete successfully before deployment. - The deployment stage packages the agent, places it in the staging bucket, and publishes it to Agent Engine. The recommended approach is to combine GitLab’s built-in security checks and Workload Identity Federation with the ADK CLI. This provides a secure, keyless, and repeatable deployment process for Google AI agents.

Read original(opens in new tab)
woowahanOriginal article

Delivering the Future: Global Hackathon (opens in new tab)

The Global Hackathon 2025 served as a massive collaborative initiative to unite over 270 technical employees from seven global entities under DeliveryHero’s umbrella, including Woowa Brothers. By leveraging the community-building expertise of the Woowahan DevRel team, the event successfully bridged geographical and technical gaps to foster innovation in "Delivering the Future." The hackathon concluded with high-level recognition from global leadership and a strategic partnership with Google Cloud, demonstrating the power of synchronized global technical synergy. ## Strategic Planning and Global Coordination * The event adopted a hybrid "Base Camp" model, where participants worked from their local entity offices while staying connected through 24-hour live streaming and centralized online channels. * Organizers meticulously navigated the logistical hurdles of spanning 70 countries, including coordinating across vastly different time zones and respecting local public holidays and vacation seasons. * Efficiency was maintained through a decentralized communication strategy, using entity-specific meetings and comprehensive guidebooks rather than frequent global meetings to prevent "meeting fatigue" across time zones. ## Technical Infrastructure and Regulatory Compliance * To accommodate diverse technical preferences, the infrastructure had to support various stacks, including AWS, Google Cloud Platform (GCP), and specific machine learning models. * The central organization team addressed complex regulatory challenges, ensuring all sandbox environments complied with strict global security standards and GDPR (EU General Data Protection Regulation). * A strategic partnership with Google Cloud provided a standardized Google AI-based environment, enabling teams to experiment rapidly with mature tools and cloud-native services. ## Local Operations and Cross-Entity Collaboration * Physical office spaces were transformed into immersive hackathon hubs to maintain the high-intensity atmosphere characteristic of offline coding marathons. * The event encouraged "office sharing" between entities located in the same city and even supported travel for members to join different regional base camps, fostering a truly global networking culture. * Local supporters used standardized checklists and operational frameworks to ensure a consistent experience for participants, whether they were in Seoul, Berlin, or Dubai. Building a successful global technical event requires a delicate balance between centralized infrastructure and local autonomy. For organizations operating across multiple regions, investing in shared technical sandboxes and robust communication frameworks is essential for turning fragmented local talent into a unified global innovation engine.

googleOriginal article

Teaching Gemini to spot exploding stars with just a few examples (opens in new tab)

Researchers have demonstrated that Google’s Gemini model can classify cosmic events with 93% accuracy, rivaling specialized machine learning models while providing human-readable explanations. By utilizing few-shot learning with only 15 examples per survey, the model addresses the "black box" limitation of traditional convolutional neural networks used in astronomy. This approach enables scientists to efficiently process the millions of alerts generated by modern telescopes while maintaining a transparent and interactive reasoning process. ## Bottlenecks in Modern Transient Astronomy * Telescopes like the Vera C. Rubin Observatory are expected to generate up to 10 million alerts per night, making manual verification impossible. * The vast majority of these alerts are "bogus" signals caused by satellite trails, cosmic rays, or instrumental artifacts rather than real supernovae. * Existing specialized models often provide binary "real" or "bogus" labels without context, forcing astronomers to either blindly trust the output or spend hours on manual verification. ## Multimodal Few-Shot Learning for Classification * The research utilized few-shot learning, providing Gemini with only 15 annotated examples for three major surveys: Pan-STARRS, MeerLICHT, and ATLAS. * Input data consisted of image triplets—a "new" alert image, a "reference" image of the same sky patch, and a "difference" image—each 100x100 pixels in size. * The model successfully generalized across different telescopes with varying pixel scales, ranging from 0.25" per pixel for Pan-STARRS to 1.8" per pixel for ATLAS. * Beyond simple labels, Gemini generates a textual description of observed features and an interest score to help astronomers prioritize follow-up observations. ## Expert Validation and Self-Assessment * A panel of 12 professional astronomers evaluated the model using a 0–5 coherence rubric, confirming that Gemini’s logic aligned with expert reasoning. * The study found that Gemini can effectively assess its own uncertainty; low self-assigned "coherence scores" were strong indicators of likely classification errors. * This ability to flag its own potential mistakes allows the model to act as a reliable partner, alerting scientists when a specific case requires human intervention. The transition from "black box" classifiers to interpretable AI assistants allows the astronomical community to scale with the data flood of next-generation telescopes. By combining high-accuracy classification with transparent reasoning, researchers can maintain scientific rigor while processing millions of cosmic events in real time.

datadog3 min readCurated summary

2023-03-08 incident: A deep dive into the platform-level recovery

Datadog’s March 8, 2023 outage removed 60% of its compute capacity, forcing teams to restore infrastructure in stages while accounting for regional and cloud-provider differences. In EU1, recovery depended on rebooting affected nodes, restoring Kubernetes control planes in a strict hierarchy, and gradually bringing application capacity back online. Scaling afterward exposed infrastructure limits that had not been considered during normal operations. ## EU1 Platform Recovery - A system patch disconnected affected EU1 nodes from the network, but the nodes could be recovered through reboots. - Recovery was initially slowed by the lack of observability and unavailable Kubernetes APIs. - Datadog operates: - **Parent clusters**, which host the control-plane pods for other clusters. - **Child clusters**, where Datadog applications run. - This hierarchy allows Datadog to use Kubernetes deployment, replacement, rolling-update, and autoscaling capabilities for child-cluster control planes. - Parent-cluster control planes run on VMs and are managed with `systemd`. ## Restoring Kubernetes Clusters Because both parent and child environments were affected by the Ubuntu 22.04 issue, recovery had to follow a strict sequence: - **Parent control planes:** Nodes running Cilium were rebooted to restore network connectivity. This finished by 08:45 UTC. - **Child control planes:** All parent-cluster nodes hosting child control-plane pods were rebooted. This finished by 09:30 UTC. - **Application nodes:** Thousands of instances across dozens of child clusters were restarted. - Recovery reached 60% by 10:20 UTC. - All application nodes were restored by 12:05 UTC. - Restarts were prioritized by workload importance and paced to avoid overwhelming Kubernetes control planes. ## Scaling Capacity and Recovering Backlogs After restoring the clusters, Datadog needed substantial additional capacity to process data buffered during the outage. - EU1 hit a Google Cloud mesh limit of **15,500 VM instances** at 14:18 UTC. - Instance creation failures became apparent around 15:00 UTC. - Datadog had not checked this documented limit before the incident, but Google Cloud quickly raised it after Datadog submitted a high-priority request. - Autoscaling also exhausted the IP capacity of subnets used by three log- and trace-processing clusters. - These clusters normally used about 35–45% of their IP capacity, but the backlog caused autoscaling to request more than twice their usual replica counts, filling the subnets. ## Practical Lessons The recovery demonstrated that restoring compute capacity is not enough: teams must also understand dependency order, control-plane architecture, cloud-provider quotas, and network-address limits. Capacity planning should account for severe backlog-driven scaling, not just normal operating utilization, and documented infrastructure limits should be validated before emergencies occur.

Read original(opens in new tab)