Augmented Reality

7 posts

figma3 min readCurated summary

What Does the Future of Software Look Like? | Figma Blog

AI may reshape software around more human, contextual interactions rather than fixed menus and mechanical commands. The post argues that future interfaces could understand intent through voice, gesture, emotion, and situation, adapting their behavior to each person. Instead of forcing users to adapt to increasingly powerful systems, software could meet users where they are while encouraging focus, presence, and healthier technology habits. ## Ephemeral Tools - Controls appear only when users select an object and indicate what they want to do. - Contextual options replace persistent menus, panels, and modes. - A video editor, for example, might show timing, pacing, alternate cuts, and sound options around a selected clip. - This lets creators focus on decisions and intent rather than remembering how software is organized. ## Magic Marker - Users interact through a combination of voice, cursor movement, gestures, sound effects, and body language. - Someone could circle an object, drag it into position, and verbally request a change. - AI would interpret these signals together, making it feel more like collaborating with a teammate. - This reduces the need for precise prompt engineering or complex document references. ## Adaptive Presence - Intelligent systems adjust their communication style and level of assistance based on user behavior. - They might offer structured guidance when someone is confused, step back when help is unnecessary, or switch between text, voice, and visuals. - Software could change pacing, simplify language, and divide information into smaller steps. - This approach is especially valuable in healthcare and education, where differences in user readiness can have serious consequences. ## Empathetic Flows - Interfaces could infer emotional states from typing speed, stylus pressure, speech patterns, facial expressions, and repeated revisions. - A food app might reduce choices when someone appears overwhelmed. - A creative tool could become quiet when the user is concentrating, while a hotel app might stop promoting upgrades when the guest seems tired. - Rather than requiring users to explicitly state what they need, systems would respond to behavioral signals. ## Situational Cues - Sound, motion, pacing, progress indicators, and visual transitions can help users understand where they are in an experience. - Earlier digital products used cues such as dial-up sounds, progress bars, and “You’ve got mail” announcements to provide orientation. - Future interfaces should counteract the overstimulation caused by attention-driven notifications. - Persistent progress indicators, transition sounds, and consistent visual language could help users regulate their attention and nervous systems. ## Spatial Tuning - Users could control software through bodily movement instead of conventional tapping and clicking. - Examples include shaping music with hand movements, navigating augmented reality by changing body orientation, or adjusting design elements through gestures. - These interactions demand attention and presence, making them harder to rush or automate. - Technology becomes an experience that intentionally slows users down rather than continually rewarding speed. ## Mash-Ups - Future systems could combine any two inputs—files, objects, sounds, locations, or physical gestures—to create something new. - The system would synthesize the combined inputs while blending their structure, tone, and meaning. - Possible examples include merging a playlist with a city map or combining digital objects through touch or gestures. The overall recommendation is to design AI-powered software around human intent, context, emotion, and physical presence. The most successful future interfaces may be those that make technology feel less like a collection of controls and more like an adaptable, considerate collaborator.

Read original(opens in new tab)
figma2 min readCurated summary

Press Start: How Controllers Shaped Video Game Design—and Where Interfaces May Go Next | Figma Blog

Video game interfaces evolved from simple start screens into complex systems of menus, maps, inventories, and customization. Yet the controller’s directional inputs and buttons have remained a consistent foundation, shaping navigation patterns and preserving players’ muscle memory. The article argues that understanding how input devices define interfaces will be essential as AR, VR, and other technologies introduce new ways to interact with digital worlds. ## From Simple Screens to Complex Game Hubs - Early games such as *Tennis for Two*, *Pong*, *PAC-MAN*, and *Super Mario Bros.* offered minimal interfaces, often requiring only a coin insertion or a start button. - Modern AAA games support many features, including: - Multiple game modes - Character customization - Cosmetics stores - Maps and inventories - Quests and settings - These features have transformed games into interconnected hubs of screens, menus, and tabs. - The growth of gaming has helped make it a major cultural and economic industry, supported by livestreaming, esports, and Olympic recognition. ## Controllers as a Design Foundation - From the Atari 2600 to the PlayStation 5, controllers have consistently combined directional controls with buttons. - This consistency allows players to transfer familiar muscle memory between generations of games. - Established controller conventions create predictable navigation patterns for players. - Designers and developers benefit from these conventions because they reduce the need to create entirely new interaction models for every game. ## The Input Defines the Interface - Every platform has a dominant input method that shapes its interface: - Desktop: mouse - Mobile: fingers and touch gestures - Television: remote controls - Mouse interfaces can use small, dense controls because the pointer is fast and precise. - Touch interfaces require larger, more widely spaced controls because fingers are less precise than cursors. - Mobile design also takes advantage of gestures such as tapping, swiping, and pinching. - The article’s broader point is that interface design must reflect the physical capabilities and limitations of the device used to navigate it. ## Preparing for New Forms of Interaction - Game interfaces increasingly connect 2D navigation systems with 3D worlds. - AR and VR controllers may allow players to engage with their physical surroundings as well as digital environments. - As input methods evolve, designers will need to rethink how people move through interfaces, rather than simply adapting existing controller-based patterns. The practical recommendation is to treat input as a core design constraint. By studying how controllers shaped game navigation—and how emerging technologies change physical interaction—designers can build interfaces that feel intuitive across future platforms.

Read original(opens in new tab)
metaOriginal article

How We Built Meta Ray-Ban Display: From Zero to Polish (opens in new tab)

Meta's development of the Ray-Ban Display AI glasses focuses on bridging the gap between sophisticated hardware engineering and intuitive user interfaces. By pairing the glasses with a neural wristband, the team addresses the fundamental challenge of creating a high-performance wearable that remains comfortable and socially acceptable for daily use. The project underscores the necessity of iterative refinement and cross-disciplinary expertise to transition from a technical prototype to a polished consumer product. ### Hardware Engineering and Physics * The design process draws parallels between hardware architecture and particle physics, emphasizing the high-precision requirements of miniaturizing components. * Engineers must manage the strict physical constraints of the Ray-Ban form factor while integrating advanced AI processing and thermal management. * The development culture prioritizes the celebration of incremental technical wins to maintain momentum during the long cycle from "zero to polish." ### Display Technology and UI Evolution * The glasses utilize a unique display system designed to provide visual overlays without obstructing the wearer’s natural field of vision. * The team is developing emerging UI patterns specifically for head-mounted displays, moving away from traditional touch-screen paradigms toward more contextual interactions. * Refining the user experience involves balancing the information density of the display with the need for a non-intrusive, "heads-up" interface. ### The Role of Neural Interfaces * The Ray-Ban Display is packaged with the Meta Neural Band, an electromyography (EMG) wristband that translates motor nerve signals into digital commands. * This wrist-based input mechanism provides a discrete and low-friction way to control the glasses' interface without the need for voice commands or physical buttons. * Integrating EMG technology represents a shift toward human-computer interfaces that are intended to feel like an extension of the user's own body. To successfully build the next generation of wearables, engineering teams should look toward multi-modal input systems—combining visual displays with neural interfaces—to solve the ergonomic and social challenges of hands-free computing.

googleOriginal article

Sensible Agent: A framework for unobtrusive interaction with proactive AR agents (opens in new tab)

Sensible Agent is a research prototype designed to move AR agents beyond explicit voice commands toward proactive, context-aware assistance. By leveraging real-time multimodal sensing of a user's environment and physical state, the framework ensures digital help is delivered unobtrusively through the most appropriate interaction modalities. This approach fundamentally reshapes human-computer interaction by anticipating user needs while minimizing cognitive and social disruption. ## Contextual Understanding via Multimodal Parsing The framework begins by analyzing the user's immediate surroundings to establish a baseline for assistance. * A Vision-Language Model (VLM) processes egocentric camera feeds from the AR headset to identify high-level activities and locations. * YAMNet, a pre-trained audio event classifier, monitors environmental noise levels to determine if audio feedback is appropriate. * The system synthesizes these inputs into a parsed context that accounts for situational impairments, such as when a user’s hands are occupied. ## Reasoning with Proactive Query Generation Once the context is established, the system determines the specific type of assistance required through a sophisticated reasoning process. * The framework uses chain-of-thought (CoT) reasoning to decompose complex problems into intermediate logical steps. * Few-shot learning, guided by examples from data collection studies, helps the model decide between actions like providing translations or displaying a grocery list. * The generator outputs a structured suggestion that includes the specific action, the query format (e.g., binary choice or icons), and the presentation modality (visual, audio, or both). ## Dynamic Modality and Interaction Management The final stage of the framework manages how the agent communicates with the user and how the user can respond without breaking their current flow. * The prototype, built on Android XR and WebXR, utilizes a UI Manager to render visual panels or generate text-to-speech (TTS) prompts based on the agent's decision. * An Input Modality Manager activates the most discreet response methods available, such as head gestures (nods), hand gestures (thumbs up), or gaze tracking. * This adaptive selection ensures that if a user is in a noisy room or a social setting, the agent can switch from verbal interaction to subtle visual cues and gesture-based confirmations. By prioritizing social awareness and context-sensitivity, Sensible Agent provides a blueprint for AR systems that feel like helpful companions rather than intrusive tools. Implementing such frameworks is essential for making proactive digital assistants practical and acceptable for long-term, everyday use in public and private spaces.

figma3 min readCurated summary

Meet the Makers Defining Tech’s Next Chapter | Figma Blog

The post previews Config 2025 and introduces speakers exploring how technology can become more human-centered. Their work spans AI-native devices, creator tools, AI interfaces, and expressive robotics. Together, they suggest that technology’s next chapter will depend not only on technical capability, but also on accessibility, empathy, and thoughtful design. ## Config 2025 and Its Vision - Figma’s Config conference will take place in San Francisco from May 6–8 and in London on May 14. - The event will examine how AI and automation are changing the way people create and experience technology. - Figma emphasizes tools and experiences that respond to evolving human needs. - In-person London tickets are sold out, while San Francisco tickets and free virtual attendance remain available. ## Building the Next Computing Platform - Andrew “Boz” Bosworth, Meta’s CTO and Head of Reality Labs, began coding through a 4-H club and later taught Mark Zuckerberg’s AI class at Harvard. - As Facebook’s 10th engineer, he helped create the News Feed. - He now leads Meta’s work on AR glasses, mixed-reality headsets, and the metaverse. - Bosworth envisions AI-native devices and proactive, personalized assistants becoming part of everyday computing. - At Config, he will discuss Meta’s efforts to develop a new computing platform. ## Making Creator Tools More Inclusive - Ebi Atawodi, Director of Product Management for YouTube Studio, combines experience in literature, technology, and product leadership. - Her work focuses on helping creators develop ideas and reach new audiences, including through AI-powered features. - Atawodi stresses that technology must be accessible and designed for people who have historically felt excluded by products. - Her approach centers on creating tools that serve a broad range of creators rather than assuming a single user experience. ## Designing Interfaces for Advanced AI - Joel Lewenstein, Head of Product Design at Anthropic, works on the interface for Claude. - He compares designing AI systems to watching a child discover new abilities, highlighting the uncertainty and possibility of emerging technology. - AI’s growing sophistication challenges traditional assumptions about how software interfaces should work. - Lewenstein argues that strong designers are defined more by how they think than by experience in a particular industry. ## Giving Robots Personality - Dr. Madeline Gannon uses her research studio, Atonaton, to transform industrial robots into expressive, lifelike mechanical creatures. - Supported by a Knight Foundation grant, she investigates how art and empathy can influence relationships with autonomous machines. - Her work explores how easily people attribute life and intention to nonliving objects. - Rather than reproducing the complexity of the human mind, Gannon demonstrates how movement, appearance, and context can make robots feel animated. Config’s featured makers represent a broad view of technological progress: more capable devices and systems must also be accessible, intuitive, expressive, and emotionally resonant. The conference aims to explore that balance through the experiences of people actively shaping AI, design, and robotics.

Read original(opens in new tab)
figma2 min readCurated summary

Charmaine Lee’s 10 Rules for Building Developer Tools | Figma Blog

Developer tools feel “magical” not because of polished interfaces, but because they help users quickly become confident creators. Charmaine Lee argues that teams should optimize what happens after the initial “aha moment,” shorten the path to meaningful creation, and build products through close, authentic engagement with developers. ## Prioritize lasting adoption over onboarding - The first-time user experience (FTUE) should not be overloaded with every product capability. - The real measure of success is whether users understand and continue using the product after the initial discovery moment. - Lens Studio 5.0’s public beta omitted a formal FTUE, testing whether the product was intuitive enough to use independently. ## Shorten the path to “magic” - Teams should identify how long it takes users to move from downloading a tool to creating and sharing something valuable. - Lens Studio’s team mapped a 19-step journey from visiting the website to submitting a first project. - By removing unnecessary steps and avoiding guidance for actions users already understood, they reduced the experience to four key moments. - User-journey mapping and testing help reveal which steps create delight, friction, or confusion. ## Meet developers in their communities - Product managers should engage directly with developers at meetups, conferences, hackathons, livestreams, and online communities. - Charmaine monitors AR discussions and attends events to learn developers’ language and gather candid feedback. - Building long-term context from these conversations enables better product decisions and more informed responses to user needs. ## Replace traditional marketing with DevRel - Developers tend to respond poorly to conventional marketing and prefer authentic communication. - Effective developer relations includes: - Real experiences and detailed product-building stories - Transparency about mistakes and limitations - A balance between accessible explanations and technical depth - DevRel should be a company-wide responsibility, not limited to a specialized team. - When employees advocate for both the product and its users, they can foster a more loyal developer community. The excerpt’s central recommendation is to design for users’ sustained progress, not merely their first impression: remove unnecessary friction, understand developers firsthand, and communicate with them honestly.

Read original(opens in new tab)
figma2 min readCurated summary

Google’s AR Design Guidelines suffice while Apple’s fall short | Figma Blog

Mobile AR platforms from Apple and Google made immersive 3D experiences accessible to mainstream app developers, but their design guidance remains limited. The author argues that Google’s guidelines are more practical and comprehensive than Apple’s, particularly for object placement and onboarding, yet both focus mostly on simple single-scene experiences. Designers must therefore develop their own best practices for complex, interactive, and collaborative AR applications. ## AR Design Becomes Mainstream - ARKit and ARCore shifted 3D UX from specialized VR hardware to mobile devices potentially reaching hundreds of millions of users. - Designers accustomed to mature 2D workflows faced unfamiliar 3D concepts and fragmented tools borrowed from gaming, film, architecture, and engineering. - Apple and Google responded with dedicated augmented-reality design guidelines. ## Where Both Guidelines Fall Short - Both focus primarily on simple object placement and sticker-like experiences. - They do not adequately address: - Object selection and hidden-object interaction - Conditional behaviors - Branching storyboards and scene flows - Teleportation and portals - Physical gestures - Multi-scene applications - Triggered or timed animation - Shared and collaborative AR environments - This narrow scope limits guidance for creating deeper personalization, state changes, and immersive interaction. ## Google’s Guidelines Compared with Apple’s - Apple’s guidance is relatively brief and concentrates largely on placing objects on detected surfaces. - Google provides more practical coverage, including: - Tap-to-place - Drag-to-place - Free placement - Realism and basic physics - UI components - Onboarding and experience design - The author questions Google’s assumptions that objects should generally be placed within the user’s reach, since AR may also involve throwing, pointing, or interacting with distant objects. - Despite its gaps, Google’s guidance is considered the stronger starting point and appears to be evolving. ## Designers Must Define New Best Practices - Official guidelines cannot keep pace with the breadth of experimentation by AR designers and developers. - Individuals such as Unity designer Bushra Mahmood have created more extensive independent AR guidance. - Other teams are assembling new workflows from existing tools to prototype functional AR applications. - This grassroots experimentation may help establish 3D UX conventions before immersive headsets become widespread. Apple and Google provide useful foundations, but neither adequately supports complex AR design. Designers should use Google’s more developed guidance as a starting point while actively creating and sharing patterns for multi-scene, interactive, animated, and collaborative experiences.

Read original(opens in new tab)