AI is reshaping design by blurring the boundary between code and canvas, while expanding—not eliminating—the need for designers. Figma’s research suggests designers are adapting to new expectations by strengthening both AI-related capabilities and enduring creative fundamentals. The future favors people who can move fluidly across tools, teams, and stages of product development.
## AI’s impact on design work
- 91% of surveyed designers say AI tools are helping them improve their work.
- “Better design” means different things to different designers, including:
- Visual polish
- More thoughtful problem-solving
- More intuitive user experiences
- These differing priorities influence how designers understand and experience their jobs.
- Design is increasingly defined by outcomes and problem-solving rather than by a single medium.
## Design hiring remains strong
- AI is not reducing demand for designers according to Figma’s research.
- 82% of surveyed hiring managers say their need for designers has either remained stable or increased.
- Demand is growing beyond technology companies.
- Organizations are seeking designers who can help translate new AI capabilities into useful products and experiences.
## Skills for the AI era
- Designers are exploring emerging practices such as:
- Prompting
- MCP-related workflows
- Connecting AI tools and processes
- Translating between design, engineering, product, and other teams
- AI-specific skills complement rather than replace foundational design abilities.
- Communication, judgment, craft, and the ability to understand user and business needs remain essential.
- The strongest designers are likely to combine technical fluency with human-centered thinking.
## Product teams are prototyping earlier
- Product managers are using Figma Make to explore ideas and build conviction more quickly.
- Teams at ServiceNow, Ticketmaster, and Affirm use prototypes to:
- Communicate complex product behaviors
- Test and develop ideas
- Make better roadmap decisions
- Prototyping is becoming accessible beyond traditional design roles.
## Code and canvas converge
- Ideas can begin in code, visual design, or anywhere in between.
- Figma presents the future of design as a continuous movement between code and canvas.
- This shift makes designers less defined by their tools and more by their ability to shape ideas across mediums.
Designers should treat AI as an extension of their creative and problem-solving toolkit, while continuing to develop core design judgment, communication, and craft. The most valuable practitioners will be those who can connect AI-enabled workflows with strong product thinking and cross-functional collaboration.
Kakao argues that monitoring hundreds of millions of daily security events cannot scale through human analysts and increasingly complex rules alone. Its solution is a hybrid AI pipeline that filters noise early, analyzes only high-value events with multiple models, and continuously improves through verified feedback. The goal is not to generate more alerts, but to understand context and identify threats worth investigating.
## The Scale Problem: Finding Threats in a Haystack
- Endpoint activity such as process execution, network connections, file changes, and privilege escalation produces hundreds of millions of events.
- The volume grows rapidly as services expand, while the proportion of genuine attacks remains very small.
- Increasing the number of analysts alongside event volume is economically and operationally unsustainable.
- AI is needed to correlate events, interpret behavior statistically and contextually, and dynamically distinguish normal activity from anomalies.
## Limitations of Rule-Based Monitoring
- Rules can identify what happened, but not why, who initiated it, or whether it fits the environment.
- Legitimate deployment commands can resemble backdoor installation, causing high false-positive rates.
- Analysis quality varies by analyst experience, shift, and time of day.
- Analysts must manually assemble host information, network sessions, process histories, and related logs into an incident narrative.
- Expanding detection categories—behavior sequences, statistical anomalies, multi-source correlations, and rare events—makes manual rule maintenance impractical.
- SIEM correlation improves on single-event rules but remains limited to predefined scenarios and struggles with unknown attack patterns.
- As rule sets and event volumes grow, both maintenance costs and matching performance become problematic.
## A Funnel-Based Hybrid Architecture
- Kakao filters events through multiple stages before using AI:
- Rule-based filters remove obvious noise.
- Learned normal patterns are automatically excluded.
- AI performs detailed analysis only on the small remainder requiring judgment.
- Rules handle clear, deterministic patterns quickly, while AI evaluates complex contextual situations.
- The framework is designed to accommodate new threat types and detection categories without creating a separate system for each scenario.
## Multi-Model Verification and Operational Resilience
- Multiple AI models independently analyze the same event and cross-check one another.
- Disagreement is treated as an uncertainty signal that can trigger deeper analyst review.
- Model diversity helps reduce bias, false positives, and missed detections.
- It also provides resilience against model failures, API outages, and quality changes after model updates.
- The design balances cost, processing speed, and accuracy rather than optimizing only for detection precision.
## Teaching AI the Environment’s Context
- Generic LLMs initially misclassified legitimate activity because they lacked knowledge of Kakao’s infrastructure.
- The system supplies structured context, including:
- Host roles
- Services running on each host
- Accounts used for automation
- Normal communication and operational patterns
- This context allows the model to act more like an analyst familiar with the organization than a generic security classifier.
## Analyzing Complete Behavior Flows
- Individual commands such as `curl`, `chmod`, and script execution can occur in both normal deployments and attacks.
- Kakao therefore reconstructs activity at the host level, linking:
- Process execution history
- Network sessions
- File changes
- Temporal ordering
- The same command can have different meanings depending on when, where, and in what sequence it occurred.
- AI evaluates the complete sequence to distinguish routine operations from intrusion behavior.
## Translating Events into AI-Usable Data
- Sending raw events directly to an LLM wastes tokens on irrelevant information and reduces accuracy.
- Different detection tasks require different signals; statistical anomaly detection and sequence analysis cannot rely on one fixed format.
- Kakao introduced:
- A standardized event schema
- Dynamic feature construction tailored to each detection type
- This reduces token usage while improving the relevance and precision of AI analysis.
## WALT: A Self-Learning Detection Loop
- Initially, analysts had to manually convert AI conclusions into new detection policies.
- Kakao developed WALT, or **Whitelist-Assisted Learning and Tuning**, to automate this feedback process.
- Repeatedly verified normal patterns are converted into exception policies.
- Those policies filter future matching events before they reach the AI engine.
- Thousands of detection policies are reportedly being generated and operated this way, allowing accuracy to improve over time.
## Cost and Performance Constraints
- Sending every event to an AI model caused unsustainable costs and processing delays.
- The funnel architecture addresses this by reserving expensive AI analysis for events that survive earlier filtering.
- The overall system must continuously balance economic cost, response speed, detection accuracy, and reliability.
Kakao’s practical recommendation is to treat AI as part of a carefully designed security pipeline—not as a replacement for rules or analysts. Effective large-scale monitoring combines deterministic filtering, contextual multi-model analysis, structured data, and a controlled feedback loop that learns from verified outcomes.
Updating security-sensitive APIs across a massive mobile codebase is difficult because vulnerable patterns may appear across hundreds of call sites and millions of lines of code. Meta’s Product Security team addresses this through secure-by-default Android frameworks and generative AI that automates migrations to those frameworks. The approach enables security patches to be proposed, validated, and submitted with minimal effort from code owners.
## Secure-by-Default Mobile Frameworks
- Meta wraps potentially unsafe Android OS APIs in frameworks designed to make secure implementations the easiest option.
- Developers are guided toward safer behavior by default rather than being expected to recognize and avoid every security risk manually.
- This strategy helps prevent a single vulnerability class from recurring across Meta’s many mobile applications.
## AI-Assisted Code Migration
- Generative AI is used to migrate existing code from unsafe APIs to the new secure frameworks.
- The system operates across millions of lines of code and numerous call sites.
- It can propose security changes, validate them, and submit patches for review.
- This reduces the manual work required from the engineers responsible for each application or codebase.
## Security at Massive Scale
- Meta’s scale—thousands of engineers, multiple apps, and billions of users—makes conventional security updates difficult to coordinate.
- The initiative combines framework design, automation, and engineering ownership to reduce friction while maintaining validation.
- The accompanying Meta Tech Podcast episode features Product Security engineers Alex and Tanu discussing the challenges and lessons from this effort.
Meta’s approach demonstrates that large-scale mobile security improvements are most practical when safer APIs and automated migration tools work together, allowing secure changes to spread broadly without requiring every engineer to perform the migration manually.
Amazon S3 began in 2006 as a simple web service for storing and retrieving objects, but its emphasis on security, durability, availability, performance, and elasticity enabled it to become foundational infrastructure. Over two decades, it scaled from roughly one petabyte to hundreds of exabytes while preserving API compatibility, reducing prices, and expanding beyond object storage. Amazon’s long-term vision is for S3 to serve as a universal foundation for data, analytics, and AI workloads.
## The Original S3 Philosophy
- S3 introduced two basic operations:
- `PUT` to store an object
- `GET` to retrieve it
- The service abstracted away complex infrastructure so developers could focus on applications.
- Its five enduring design principles are:
- **Security:** Data is protected by default.
- **Durability:** Designed for 11 nines of durability, with a lossless operating model.
- **Availability:** Failure is assumed and handled throughout the system.
- **Performance:** Storage capacity can grow without degrading performance.
- **Elasticity:** Capacity expands and contracts automatically.
## From One Petabyte to Hundreds of Exabytes
- At launch, S3 had approximately:
- One petabyte of capacity
- 400 storage nodes across 15 racks and three data centers
- 15 Gbps of bandwidth
- A maximum object size of 5 GB
- A price of $0.15 per GB
- Today, S3:
- Stores more than 500 trillion objects.
- Serves over 200 million requests per second.
- Operates across 123 Availability Zones in 39 AWS Regions.
- Supports objects up to 50 TB—10,000 times larger than the original limit.
- Storage prices have fallen by roughly 85%, to slightly above 2 cents per GB.
- S3 Intelligent-Tiering has saved customers more than $6 billion in storage costs.
- The S3 API has become an industry standard, with many other storage systems offering compatible interfaces.
## Backward Compatibility and Long-Term Reliability
- Code written against S3 in 2006 still works without modification.
- AWS has repeatedly replaced disks, storage systems, and request-processing code while preserving access to older data.
- This compatibility reflects S3’s goal of remaining infrastructure that “just works” despite continuous internal change.
## Engineering for Durability and Scale
- Microservices continuously inspect every byte across the fleet.
- Auditor services detect degradation and automatically trigger repair and re-replication.
- Automated formal methods mathematically verify correctness in areas such as:
- The index subsystem
- Cross-Region replication
- Access policies
- AWS has progressively rewritten performance-critical components in Rust over the past eight years.
- Rust improves performance while preventing memory-safety bugs and other classes of errors at compile time.
- S3 follows the principle that scale should improve the service: larger, more distributed workloads become increasingly decorrelated, improving reliability for all customers.
## S3 as a Foundation for Data and AI
Amazon’s future vision is for customers to store data once in S3 and work with it directly, avoiding costly copies and specialized systems.
- **S3 Tables** provides managed Apache Iceberg tables with automated maintenance to improve query performance and reduce storage costs.
- **S3 Vectors** supports semantic search and retrieval-augmented generation, with up to 2 billion vectors per index and sub-100 ms query latency.
- Within five months of launch, customers created over 250,000 indexes, ingested more than 40 billion vectors, and executed over 1 billion queries.
- **S3 Metadata** enables centralized, faster data discovery without recursively listing large buckets.
These additions extend S3 from inexpensive object storage into a broader platform for analytics, search, and AI while retaining its scale and cost advantages.
LY Corporation is consolidating the former LINE “Verda” and Yahoo Japan “YNW” private clouds into Flava, a next-generation platform designed for large-scale, uninterrupted operations. Its approach assumes failures will occur, prioritizing stateless services, application-led availability, rapid IaC-based recovery, and extensive automation. Flava also restructures the architecture around shared resources, upstream OpenStack, default VPC networking, and user-driven cost optimization.
## Failure-Aware Design and Operations
- VM root disks are treated as temporary; persistent data is placed in external storage so instance failures have limited service impact.
- Availability is achieved through cooperation between infrastructure and applications rather than excessive infrastructure-side guarantees.
- Recovery focuses on maintaining service continuity, rebuilding environments quickly with infrastructure as code, and avoiding lengthy root-cause investigations during incidents.
- The company promotes KaaS and PaaS to help developers build resilient services without managing low-level infrastructure.
- OS configuration, package installation, networking, and other changes are managed as code through CI/CD.
- Deployments are performed by availability zone to limit the blast radius of failures.
## Observability from Fleet-Wide Trends to Root Causes
- Prometheus, Grafana, and custom dashboards monitor overall cloud health and long-term trends.
- When anomalies appear, engineers investigate at a deeper level using kernel traces, packet captures, and other low-level diagnostics.
- This combination of broad monitoring and detailed investigation allows teams to move between “forest” and “tree” perspectives.
- The operational model depends not only on tools but also on engineers capable of tracing problems down to their fundamental causes.
## OSS, Software-Defined Infrastructure, and Custom Development
- The platform relies heavily on OpenStack, Envoy, Linux kernel technologies such as eBPF/XDP, FRR, and Ceph.
- LY contributes patches and new capabilities upstream instead of maintaining long-lived private forks.
- It has developed SRv6 BGP functionality required for Flava’s VPCs and contributed related work to FRRouting and the Linux kernel.
- Compute, VPC, DNS, and load-balancing services run primarily on commodity x86 servers rather than specialized appliances.
- XDP-based data planes, hardware offload, and system tuning are used to achieve near-wire-speed throughput and low latency.
- Where OSS cannot meet internal requirements, LY builds systems from scratch, including the Dragon object store, SDN control-plane components, load-balancer health agents, and service discovery tools written in Rust, Go, and Python.
## Autonomous Hardware Operations
- With tens of thousands of hypervisors and petabyte-scale storage, hardware failures occur continuously.
- Failure detection, requests to data-center technicians, hardware replacement, and cluster reintegration are largely automated.
- Some exceptional cases still require engineers, but LY plans to use LLMs to automate more of these operational tasks.
## Flava’s Architectural Improvements
### Shared Resource Pools
- Older clouds used many dedicated clusters and resource pools, making capacity planning complex and reducing utilization.
- Flava consolidates most products and services into one large shared resource pool.
- This reduces planning variables, improves resource efficiency, and accelerates provisioning.
### Upstream-Compatible OpenStack
- Excessive customization in the legacy environment made upgrades difficult.
- Flava minimizes private patches, follows upstream OpenStack, and contributes necessary improvements back to the project.
- This enables regular upgrade cycles and keeps security fixes and features current.
### VPC by Default
- VPC networking is the standard security model for multi-tenant workloads.
- Logical isolation replaces many cases where dedicated VLANs or firewalls previously required months of preparation.
- Equivalent security environments can now be provisioned in minutes.
- The VPC data plane is being redesigned with XDP to support the reliability and performance required at company-wide scale.
### Built-In Cost Optimization
- Development environments require resource lifetimes, allowing unused “zombie” resources to be deleted automatically.
- Object storage offers bucket classes such as “High Performance” and “Scalable.”
- Users can change storage classes without changing endpoints, adapting cost and performance as access patterns evolve.
## Remaining Challenges
- Flava currently offers only a limited set of products and must expand its capabilities while addressing post-launch bugs and overlooked requirements.
- The largest challenge is migrating users from the legacy platforms.
- LY is working to provide transparent migration tools and reduce manual effort while shortening the period of duplicate investment in old and new infrastructure.
## Team and Engineering Culture
- The team includes specialists ranging from kernel developers to web-front-end engineers.
- Engineers are expected to understand and control infrastructure rather than treat it as a black box.
- Deep source-level expertise enables upstream OSS contributions and informed negotiations with commercial vendors.
- This culture of ownership and technical control is presented as a core reason the platform can evolve at LY’s scale.
LY’s experience demonstrates that large private clouds can combine OSS, custom software, commodity hardware, and rigorous automation effectively. The practical recommendation is to design for failure, keep infrastructure reproducible through IaC, contribute changes upstream where possible, and use custom development selectively for requirements that general-purpose platforms cannot satisfy.
Amazon S3 now lets customers create general purpose buckets in an account regional namespace, making bucket names predictable and reusable across AWS Regions. Names combine a customer-selected prefix with an account-, Region-, and namespace-specific suffix, preventing other accounts from claiming them. The feature preserves existing general purpose bucket capabilities while improving governance and automation.
## Account Regional Bucket Namespaces
- Bucket names use a format such as `mybucket-123456789012-us-east-1-an`.
- The suffix identifies the AWS account and Region, ensuring that other accounts cannot create buckets using it.
- The combined prefix and suffix must be between 3 and 63 characters.
- Buckets support the same features as general purpose buckets in the global namespace.
## Governance and Policy Controls
- IAM policies and AWS Organizations service control policies can enforce namespace usage.
- The new `s3:x-amz-bucket-namespace` condition key allows organizations to require account regional bucket creation.
## Creating Buckets
- In the S3 console, select **Account regional namespace** when creating a bucket.
- AWS CLI requests use the `--bucket-namespace account-regional` option.
- SDKs can pass `BucketNamespace: "account-regional"` to the `CreateBucket` API.
- Applications can use STS to retrieve the account ID and the SDK’s Region to construct compliant names.
## Infrastructure as Code
- CloudFormation templates can use `AWS::AccountId` and `AWS::Region` to construct bucket names.
- The `BucketNamespace: "account-regional"` property enables the feature.
- `BucketNamePrefix` can be used when only the customer-defined prefix should appear in the template; AWS adds the account regional suffix automatically.
## Limitations and Availability
- Existing global-namespace buckets cannot be renamed into the account regional namespace; new buckets must be created.
- The feature applies only to S3 general purpose buckets.
- S3 table and vector buckets use account-level namespaces, while directory buckets use zonal namespaces.
- It is available in 37 AWS Regions, including AWS China and GovCloud Regions, with no additional cost.
Organizations can adopt account regional namespaces to simplify bucket provisioning, prevent naming conflicts, and enforce consistent naming through IAM, Organizations policies, and infrastructure-as-code tools.
GitHub built a continuous, AI-assisted accessibility feedback system to replace scattered reports, unclear ownership, and unresolved “phase two” promises. The workflow combines GitHub Actions, Copilot, and GitHub Models to turn user feedback into tracked, prioritized issues while preserving human judgment. Its goal is continuous follow-through: every accessibility barrier is captured, routed, reviewed, and acted upon.
## Accessibility as a Living System
- GitHub treats accessibility as an ongoing methodology rather than a one-time audit or standalone product.
- The approach combines:
- Automation
- Artificial intelligence
- Human expertise
- Real user feedback is considered more valuable than automated code scans because it reveals barriers experienced in real workflows.
- The system supports GitHub’s 2025 Global Accessibility Awareness Day pledge to improve accessibility across the open source ecosystem.
- Technology helps process feedback at scale, turning unstructured reports into clearer, implementation-ready work.
## Designing for Different Users
The workflow was designed around three primary groups:
- **Issue submitters**
- Community managers, support agents, and sales representatives submit reports for users and customers.
- Since they may not be accessibility specialists, the system guides them and teaches accessibility concepts during submission.
- **Accessibility and service teams**
- Engineers and designers need actionable reports containing reproducible steps, WCAG references, severity ratings, and ownership information.
- **Program and product managers**
- Leaders need trend data, issue categories, and progress visibility to prioritize investments.
The design treats feedback as data moving through a pipeline and allows the process to evolve over time.
## Event-Driven Feedback Workflow
- Each workflow stage triggers a GitHub Action that determines what happens next.
- Key events include:
- New issues launching Copilot analysis through the GitHub Models API
- Status changes initiating hand-offs between teams
- Resolutions triggering follow-up with the original submitter
- Actions can be started manually or rerun, allowing humans to intervene whenever necessary.
- GitHub initially built the system largely by hand in mid-2024; newer tools such as Agentic Workflows could now create similar Actions from natural-language instructions.
- The workflow contains seven stages:
- Intake
- Copilot analysis
- Submitter review
- Accessibility team review
- Link audits
- Closing the loop
- Improvement
- Feedback loops allow submitters to rerun analysis, resolved issues to return for further review, and improvements to update Copilot prompts.
## Actioning Intake
- Accessibility feedback can arrive through support tickets, social media, email, direct outreach, or GitHub’s accessibility discussion board.
- Approximately 90% of feedback currently comes through the public discussion board.
- Public discussions let other users:
- Confirm reported problems
- Add context
- Share workarounds
- Reports from the community often contain more detail than conventional support tickets.
- GitHub acknowledges every report within five business days, including reports it cannot directly address.
- When internal action is needed, a team member creates a tracking issue using a custom accessibility feedback template.
- The template records:
- The user’s original report
- The feedback source
- Relevant product components
- This preserves important context as feedback moves from intake into triage.
AI-driven commerce is emerging quickly, but making it reliable requires much more than adding an AI checkout button. Sellers must manage fragmented catalog integrations, real-time inventory and variant data, evolving protocols, secure payment tokens, fraud detection, fulfillment, and post-purchase operations. The central recommendation is to use adaptable infrastructure and begin with a limited, measurable product selection rather than launching an entire catalog at once.
## Catalog Integration and Data Quality
- Product catalogs are the entry point for AI agents, but each agent may require a different format, such as:
- SFTP file drops
- Custom APIs
- Agent-specific feed specifications
- Reformatting the same catalog for multiple agents creates a costly maintenance burden.
- Reliable “ingestion-ready” data determines whether products appear consistently across AI shopping surfaces.
- A shared commerce layer can syndicate one catalog across supported agents and eliminate duplicate integrations.
## Real-Time Inventory and Product Variants
- Agents need to verify current availability immediately before presenting checkout options.
- Inventory becomes harder to manage when products include combinations of:
- Sizes
- Colors
- Customizations
- Variant-specific availability
- Checkout APIs must support real-time availability checks and alternative recommendations when a particular configuration is unavailable.
- Real-time accuracy is essential for customer trust and brand reputation.
## Protocol Evolution and Compatibility
- Agentic commerce protocols are changing rapidly, with new releases adding payment handlers, scoped tokens, discounts, buyer authentication, and transport methods.
- Sellers risk creating “zombie integrations” that become obsolete when an AI platform changes direction.
- A protocol-agnostic commerce layer can help businesses support standards such as ACP and Google’s UCP without rebuilding their systems for every change.
## Secure Payments Through Shared Payment Tokens
- Shared Payment Tokens allow agents to initiate payments with a buyer’s permission without exposing payment credentials.
- The token layer connects AI agents to existing payment rails while limiting transaction scope.
- Agentic commerce requires more than payment authorization; systems must also support:
- Product discovery
- Checkout state management
- Shipping
- Returns and refunds
- The broader infrastructure must cover the full transaction lifecycle.
## Fraud Detection Without Human Browser Signals
- Traditional fraud tools often depend on signals such as mouse movements, browser fingerprints, device details, and window size.
- Those signals disappear when an AI agent performs the transaction.
- Network-level payment history can provide risk context even when a purchase is new to a particular seller.
- Shared Payment Tokens allow fraud systems such as Radar to evaluate agentic purchases similarly to traditional checkout transactions.
- Early deployments with major retailers reportedly experienced fraud rates near zero.
## Start with a Focused Product Selection
- Sellers should avoid enabling their entire catalog immediately.
- A practical launch strategy is to:
- Select a small group of high-conversion SKUs
- Use simple products with direct-to-home fulfillment
- Monitor conversion, inventory behavior, payment methods, and fulfillment issues
- URBN initially focused on popular categories such as dresses and denim rather than its full range, which also includes complex products like plants and custom furniture.
- Early launches should function as controlled experiments that produce data for broader expansion.
## A Strategic Shift in Retail Discovery
- Agentic commerce moves buying intent from stores, websites, and branded mobile apps onto AI platforms.
- This changes how sellers must approach:
- Product discovery
- Brand control
- Trust
- Dispute resolution
- The relationship between the seller and customer
- Agents increasingly mediate product selection and purchase decisions, requiring sellers to adapt their commerce strategy beyond the traditional storefront.
Sellers should treat agentic commerce as an evolving channel rather than a one-time integration. Start with reliable data, a narrow product scope, secure tokenized payments, and infrastructure that can absorb protocol changes before scaling to more products and complex fulfillment scenarios.
The onboarding of 40 new Kakao developers shifted their perspective from making features work to designing systems that survive real-world operations. Across databases, security, and AI, they learned that there is rarely one perfect answer; the best choice depends on scale, risk, maintainability, and business needs. The central lesson was to replace theoretical correctness with responsible, adaptable engineering judgment.
## Database: From Finding the Right Answer to Preparing for Change
- Database design must be evaluated by whether it can withstand traffic, schema changes, and operational demands—not only by theoretical correctness.
- Foreign keys are not automatically the best choice:
- They can introduce locking, performance, and flexibility concerns.
- Referential integrity can instead be managed at the application layer, provided testing and correction processes are strong.
- Soft deletion, using fields such as `deleted_at`, supports auditability and recovery and is often an essential operational strategy.
- Indexes should be selected according to the questions the database must answer:
- B-tree, GIN, GiST, SP-GiST, and vector indexes serve different data and query patterns.
- Execution plans reveal whether SQL uses indexes or performs full table scans, directly affecting I/O and response times.
- Duplication is not always harmful:
- Intentional denormalization can avoid expensive joins.
- Snapshot data can simplify reads and preserve the information needed by a business workflow.
- In MongoDB, embedding selected related data can make screen queries much simpler than relying exclusively on references.
- Different database systems embody different trade-offs among performance, consistency, scalability, and operational cost.
- The training covered MySQL high availability, PostgreSQL primary-key structures, cloud-native systems such as Neon, and the broader storage-to-analysis pipeline of Hadoop and Spark.
- The resulting mindset favors designs that are safe to change and affordable to operate over designs that are theoretically perfect.
## Security and IT: From Someone Else’s Responsibility to a Personal Default
- Security became a direct consequence of developers’ code rather than merely a compliance or infrastructure concern.
- Everyday safeguards such as development/production separation, VPNs, and antivirus software demonstrate that safety often requires accepting some inconvenience.
- DDoS defense is not only about blocking traffic:
- It can be difficult to distinguish an attack from legitimate traffic spikes caused by a popular event.
- Developers should apply basic controls such as rate limiting and escalate suspicious activity through established response channels.
- Hands-on API exploitation made vulnerabilities concrete and encouraged developers to view security through an attacker’s perspective.
- Security must be continuous:
- AI is increasingly being used both to discover vulnerabilities and to strengthen attacks.
- Social-engineering methods involving QR codes, app permissions, and human behavior require more than purely technical defenses.
- Security checks should be integrated from the beginning of development, not performed only at the end.
- Software quality also depends on people:
- Code should remain understandable enough for another developer to take over quickly.
- Strong engineering means choosing and communicating the most appropriate solution for the business context, not merely finding a technically possible one.
## AI: From Chatting with Models to Designing Systems
- An AI agent is not simply a model; it is an architecture composed of tools, routing logic, error handling, and model calls.
- Agent development applies familiar software-engineering practices to probabilistic models.
- Because LLM outputs can vary, reliable systems need deliberate controls:
- Prompt chaining breaks large tasks into smaller steps and limits context contamination.
- Few-shot examples clarify required output formats.
- Routing selects different prompts or workflows based on conditions.
- Multi-agent systems divide responsibilities among specialized agents, echoing the modularity and scalability principles of microservices.
- RAG reduces hallucinations structurally by:
- Chunking documents.
- Searching for semantically similar vectors.
- Supplying retrieved information to the model as additional context.
- MCP exposes internal systems and data as callable tools, effectively enabling remote function calling and connecting AI to enterprise capabilities.
- Effective AI use shifted from criticizing poor answers to specifying clear objectives, formats, examples, context, and supporting data.
- The goal is not merely to receive an intelligent response, but to design a system that consistently produces intelligent behavior.
The training ultimately marked a transition from student-style problem solving to professional engineering. Developers should consider operational resilience, security, maintainability, and business value, then make and clearly explain the most reasonable choice for the circumstances.
The post describes Kakao’s 2026 server-engineering onboarding program, which turns uncertainty into practical understanding through structured implementation, testing, and refactoring. Rather than supplying fixed answers, the program repeatedly asks developers to explain their design decisions and assess what their tests protect. Its central lesson is that server development becomes manageable when engineers build clear reasoning, maintainable structures, and safe change processes.
## Onboarding Through Three Stages
- The program follows a progression:
1. TDD- and OOP-based implementation
2. Acceptance testing for legacy code
3. Refactoring legacy code
- The focus is not only on what to build, but on how to make engineering decisions.
- Core goals include:
- Designing maintainable structures
- Analyzing and safely improving legacy systems
- Collaborating effectively, including responsible AI usage
- Although originally designed for server developers, the program expanded to frontend, Android, and iOS engineers because engineering principles apply across technology stacks.
## Learning Through Questions and Collaboration
- Participants were repeatedly asked:
- Why was this design chosen?
- Does this object truly own this responsibility?
- What behavior does this test protect?
- Daily meetings, pair programming, troubleshooting discussions, and PR reviews made development a collaborative activity.
- The program aimed to develop engineers who could explain and defend their designs, rather than merely produce working code.
## Mission 1: Building a Lottery Game with TDD and OOP
- The first assignment implemented:
- Automatic and manual lottery purchases
- A fixed ticket price of 1,000 won
- Winning-statistics calculations
- Constraints encouraged better design:
- One level of indentation
- Methods limited to 10 lines
- Primitive values wrapped in value objects
- First-class collections
- Avoiding `else` through early returns
- TDD required tests to be written before implementation.
### Making Randomness Testable
- Random lottery-number generation initially made tests unpredictable and tightly coupled to concrete implementations.
- The solution was to:
- Introduce a number-generation interface
- Inject the generation strategy
- Create a separate test generator
- This made test results controllable and encouraged a more flexible design.
### Considering Value Objects and Caching
- The team also questioned whether identical number values should always create new objects.
- This led to discussions about caching and the difference between object identity and value equality.
- The main lesson was to evaluate design decisions, not just make the feature work.
## Mission 2: Writing Acceptance Tests for Legacy Code
- Participants first protected the existing system before modifying it.
- Tests focused on externally observable behavior:
- User actions
- System responses
- State changes
- Strong assertions verified not merely that an operation succeeded, but that it produced the correct result.
- Cucumber-based BDD expressed scenarios in a form understandable to non-developers, treating tests as shared specifications.
### Achieving Production Parity
- To avoid “works on my machine” problems, the test environment was aligned with production:
- PostgreSQL replaced H2
- Docker standardized execution environments
- Gradle tasks automated test execution
- Test-data isolation used:
- Reverse-order foreign-key deletion
- `TRUNCATE ... CASCADE`
- Shared cleanup utilities
- These measures ensured tests started from consistent, independent states.
## Mission 3: Refactoring Legacy Code Safely
- The final mission treated refactoring as training in decision-making, not simply an exercise in clean code.
- The central rule was to separate structural and behavioral changes:
- Structural changes must preserve behavior.
- Behavior changes must avoid unrelated structural modifications.
- PR reviews helped identify unintended behavior changes and taught participants to predict and control the effects of modifications.
- AI was used during refactoring to accelerate broad code changes, but large changes were difficult to verify, highlighting the need to control scope and validate changes carefully.
The onboarding’s practical recommendation is to approach server development through small, explainable decisions: write controllable tests, protect legacy behavior before changing it, separate refactoring from feature changes, and use AI as an assistant rather than a substitute for engineering judgment.
Cloudflare’s new Account Abuse Protection suite targets fraud from both bots and humans, focusing on whether activity is authentic rather than merely automated. It combines leaked-credential detection and account-takeover signals with new tools for identifying risky signups and suspicious identities. The capabilities are in Early Access for Bot Management Enterprise customers at no additional cost temporarily.
## Leaked Credentials and Account Takeover
- Cloudflare reports that 41% of network logins use leaked credentials, with password reuse allowing old breaches to compromise valuable accounts.
- Its leaked credential check compares hashed passwords against known breach data without storing or accessing plaintext passwords.
- More than 60% of login-page traffic during the 2024 Black Friday analysis was automated, enabling attackers to test stolen credentials at scale.
- Account takeover (ATO) detections identify customer-specific suspicious login behavior and expose attempted attacks in the Security analytics dashboard.
- These detections caught an average of 6.9 billion suspicious login attempts per day across Cloudflare’s network during the referenced week.
## Fraud Requires More Than Bot Detection
- Modern abuse combines automation, human fraud farms, device and location spoofing, and synthetic identities.
- Attackers may use valid credentials, operate at human speed, or employ AI agents, making simple bot classification insufficient.
- Common customer problems include fake users exploiting free trials, attackers logging in with correct passwords, and human-paced account draining.
- Effective protection must evaluate intent, identity, and authenticity alongside automation.
## Detecting Suspicious Account Creation
- Disposable email addresses allow attackers to create large numbers of accounts for promotions or other abuse without maintaining real email infrastructure.
- Cloudflare’s disposable email check provides a binary signal that customers can use in security rules.
- Organizations can block disposable addresses outright or challenge users who register with them.
- Cloudflare also introduces email-risk assessment based on suspicious email patterns and infrastructure, helping identify potentially fraudulent signups.
## Privacy-Preserving User Identification
- Hashed User IDs are per-domain identifiers created by cryptographically hashing usernames.
- They help customers correlate suspicious activity and mitigate fraudulent traffic without exposing users’ original identifiers.
- The feature is intended to identify risky account behavior while preserving end-user privacy.
Cloudflare recommends enabling leaked-credential checks and using the new signup, identity, and behavioral signals together. This layered approach is better suited to fraud campaigns that blend valid credentials, human activity, and automated tools.
Discord is introducing “Discord Official,” allowing game studios and publishers to claim their game profiles and verify their official Discord servers. Developers can customize profile information, assets, links, and platform details while improving player trust and discoverability. At launch, the program supports playable PC games listed on Steam, including Early Access titles.
## Benefits of Becoming Discord Official
- Creates a verified game profile players can trust when searching for or sharing the game.
- Lets developers customize:
- Game descriptions
- Cover, icon, and banner art
- Screenshots and other media
- Social links, supported platforms, and publisher details
- Connects the game profile to its verified Discord server.
- Gives the server higher visibility in Discord Discovery.
- Helps developers build communities that support updates, feedback, player engagement, and live-service development.
- Discord reports more than 10,000 game communities and 80 million members as of December 31, 2025.
## Eligibility Requirements
Before applying, developers must have:
- A studio-owned Discord server ready for verification.
- A playable game listed on Steam.
- A Steam store page that links to the game’s Discord server.
- A Team in the Discord Developer Portal.
- A new or existing Developer Portal application for the game.
- The Discord server owner added to the Developer Portal Team, since they must provide a verification code.
At launch, only PC games available in Early Access or Full Release are supported. The program is limited to the actual studios and publishers behind the games.
## Claiming a Game and Verifying the Server
- Open the game’s application in the Discord Developer Portal.
- Select **Game Identity** under **Games**, then choose **Claim Game**.
- Search for and select the game.
- Complete the Game Claim Verification form.
- The server owner receives a verification email and supplies its code.
- Submit the application for Discord’s review.
## Ongoing Developer Features
After approval, developers can update their game profiles to promote content updates, cosmetic releases, and new seasons. Discord also says it plans to introduce additional tools to make game development, testing, and community management easier.
Studios with eligible Steam games should prepare their Developer Portal team, application, and server ownership details before applying through Discord’s **How to Claim Your Game** documentation.
Google Research is expanding Flood Hub with urban flash flood forecasts that can provide up to 24 hours’ warning. The system addresses the lack of historical flood observations by using Gemini to extract verified events from public news reports, creating the Groundsource dataset for model training. Its global, lower-resolution approach aims to extend useful warnings to regions that lack expensive sensors and forecasting infrastructure, particularly in the Global South.
## The Need for Earlier Flash Flood Warnings
- Flash floods cause roughly 85% of flood-related deaths worldwide and kill more than 5,000 people annually.
- They often develop within six hours of intense rainfall, making rapid warnings essential.
- Even 12 hours of warning can reduce flood damage by about 60%.
- Early warning coverage remains highly unequal: fewer than half of developing countries have access to multi-hazard warning systems.
- Flood Hub previously focused mainly on slower-moving riverine floods, covering more than 2 billion people across 150 countries.
## The Data Problem: “Invisible” Floods
- River flood models can rely on stream gauges that record water levels and flow.
- Flash floods may occur far from gauges, especially in cities where rainfall, impermeable surfaces, drainage, and terrain interact unpredictably.
- Building detailed physical simulations globally would be computationally expensive.
- Historical, precisely located flash flood records are also scarce, preventing conventional supervised machine learning.
- Google’s Groundsource method uses Gemini to analyze public news reports, verify flood locations and times, and assemble a historical flash flood dataset.
## Scaling from Local Systems to Global Coverage
- Local flash flood systems can be highly accurate using rain sensors, radar, water-level monitors, and flow measurements.
- These systems are expensive to deploy and require location-specific calibration and engineering expertise.
- Broader systems such as WMO’s FFGS, ERIC, and the U.S. NWS warning system depend on high-resolution maps, radar forecasts, and skilled hydrologists.
- Those resources are often unavailable in the Global South.
- Google’s model instead uses globally available products, including NASA IMERG, NOAA CPC, ECMWF’s IFS HRES forecasts, and Google DeepMind’s medium-range weather model.
- Forecasts currently operate at a 20-by-20-kilometer resolution, constrained by the resolution of global data sources.
## The Urban Flash Flood Model
- The model estimates whether a flash flood is likely in a given area during the next 24 hours.
- It uses a recurrent neural network with a long short-term memory (LSTM) component to process meteorological time series.
- Inputs also include static geographic and human-environment factors:
- Urbanization density
- Topography
- Soil absorption rates
- The initial rollout targets urban regions, where news coverage is denser and most of the world’s population lives.
- It currently predicts impacts in areas with population densities above 100 people per square kilometer.
## Evaluation and Reported Performance
- Precision was measured against the Groundsource dataset, but raw precision likely understates actual performance because some genuine floods are never reported.
- A manual review of 100 alerts per continent found that many apparent false positives were confirmed flood events.
- Recall was also evaluated against major floods recorded by the Global Disaster Awareness and Coordination System (GDACS).
- Results indicate comparable precision and recall in regions such as South America and Southeast Asia and in wealthier countries with better instrumentation.
The approach demonstrates how AI and unstructured public information can help provide scalable flash flood warnings where conventional monitoring infrastructure is limited. Its current urban focus and 20-kilometer resolution make it a broad early-warning tool rather than a replacement for highly localized sensor networks.
Groundsource is a Google Research methodology that uses Gemini to convert global news reports into structured historical records of natural disasters. Its first dataset contains 2.6 million flash-flood events across more than 150 countries from 2000 onward, addressing major gaps in conventional flood databases. Google reports that the system can support near-global urban flash-flood forecasts up to 24 hours in advance.
## The problem: Limited historical disaster data
- Floods lack the standardized global sensor infrastructure available for hazards such as earthquakes.
- Existing sources, including the Global Flood Database and Dartmouth Flood Observatory, are limited by cloud cover, satellite revisit times, and their focus on large or long-lasting floods.
- GDACS contains roughly 10,000 high-impact disaster records but misses many localized and rapidly developing flash floods.
- This shortage of reliable historical data makes global forecasting, model training, and validation difficult.
## How Groundsource processes news
- The system analyzes news articles where flooding is the primary subject.
- Google Read Aloud extracts article text in 80 languages, which is translated into English using Cloud Translation.
- Gemini then applies a verification-oriented prompt to:
- Distinguish actual past or ongoing floods from warnings, policy discussions, and general risk reports.
- Resolve relative dates such as “last Tuesday” using the article’s publication date.
- Identify precise locations, including neighborhoods and streets.
- Map locations to standardized geographic polygons through Google Maps Platform.
## Accuracy and scale
- Manual evaluation found:
- 60% of events were accurate in both timing and location.
- 82% were sufficiently accurate for practical analysis, such as identifying the correct administrative district or event day.
- The resulting dataset contains 2.6 million flood events, greatly exceeding traditional monitoring archives.
- Between 2020 and 2026, Groundsource captured 85%–100% of severe flood events listed by GDACS while also recording smaller local incidents.
- Coverage is densest in recent years, particularly from 2020 to 2025, reflecting the growth of digitized news.
## Forecasting and future applications
- Groundsource data has enabled near-global urban flash-flood forecasts up to 24 hours ahead.
- These forecasts are being integrated into Google Flood Hub.
- Google plans to improve rural coverage and incorporate additional data sources.
- The same approach could help build historical datasets for droughts, landslides, avalanches, and other hazards with limited ground-truth records.
Groundsource demonstrates that news archives can serve as a large-scale source of disaster history when combined with language models, translation, and geographic verification. Its open flash-flood dataset could improve forecasting and resilience planning, though its reported accuracy levels make continued validation and refinement important.
Agentic development is reshaping software engineering at Spotify and Anthropic, from how developers write code to how organizations manage delivery. The discussion highlights Claude-powered agents, enterprise-scale context management, and the need to rethink testing, review, and accountability. The speakers conclude that agents will soon handle more of the full software lifecycle, including maintenance and deletion.
## The Opus 4.5 Inflection Point
- Spotify observed a sharp increase in agent-driven development after Opus 4.5 went online on November 25, 2025.
- Engineers increasingly shifted from working primarily in IDEs to using terminals and agent-based workflows.
- The change was presented as a practical transformation in daily engineering work, not merely an experimental trend.
## Honk: Spotify’s Background Coding Agent
- Spotify employees can invoke Honk by mentioning it in Slack.
- Honk evolved from deterministic code migrations into a Slack-native agent capable of complex migrations across thousands of repositories.
- Teams can discuss a problem in Slack and ask Honk to investigate or implement a solution directly.
- Spotify is continuing to explore how background coding agents can operate at larger scale.
## Context Engineering and Control
- Scaling agents across many repositories requires consistent, reproducible configuration.
- Anthropic recommends well-structured `CLAUDE.md` files and reusable skills that describe engineering roles, domains, and expected workflows.
- The emphasis is on simple, standardized context rather than overly complex orchestration.
- Both companies are still identifying gaps in how agents receive context and how their actions are coordinated across enterprise systems.
## Testing, Reviews, and Accountability
- Agent-generated code can be produced faster than humans can review it, creating new bottlenecks.
- Organizations must reconsider testing, governance, and approval processes as output volume increases.
- Accountability should remain tied to the outcome, regardless of whether code was produced by a human or an agent.
- The discussion frames agent adoption as an organizational change, not just a tooling upgrade.
## The Next Stage of Agentic Development
- The current phase has focused largely on code creation; the next phase will expand into maintenance, deletion, and other less popular but essential engineering work.
- Spotify is evolving Backstage from a human-oriented developer portal into an agent-first platform.
- MCP connections are expected to replace more manual developer workflows.
- Anthropic’s internal “ant-fooding” practice continues to generate product ideas from employees using its own tools, including Claude Code and Cowork.
Organizations adopting agentic development should start with reliable feedback loops, standardized context, and clear human accountability. The most significant gains will come when agents are integrated across the entire software lifecycle rather than used only for writing new code.