Prompt Engineering

35 posts

googleOriginal article

Benchmarking LLMs for global health (opens in new tab)

Google Research has introduced a benchmarking pipeline and a dataset of over 11,000 synthetic personas to evaluate how Large Language Models (LLMs) handle tropical and infectious diseases (TRINDs). While LLMs excel at standard medical exams like the USMLE, this study reveals significant performance gaps when models encounter the regional context shifts and localized health data common in low-resource settings. The research concludes that integrating specific environmental context and advanced reasoning techniques is essential for making LLMs reliable decision-support tools for global health. ## Development of the TRINDs Synthetic Dataset * Researchers created a dataset of 11,000+ personas covering 50 tropical and infectious diseases to address the lack of rigorous evaluation data for out-of-distribution medical tasks. * The process began with "seed" templates based on factual data from the WHO, CDC, and PAHO, which were then reviewed by clinicians for clinical relevance. * The dataset was expanded using LLM prompting to include diverse demographic, clinical, and consumer-focused augmentations. * To test linguistic distribution shifts, the seed set was manually translated into French to evaluate how language changes impact diagnostic accuracy. ## Identifying Critical Performance Drivers * Evaluations of Gemini 1.5 models showed that accuracy on TRINDs is lower than reported performance on standard U.S. medical benchmarks, indicating a struggle with "out-of-distribution" disease types. * Contextual information is the primary driver of accuracy; the highest performance was achieved only when specific symptoms were combined with location and risk factors. * The study found that symptoms alone are often insufficient for an accurate diagnosis, emphasizing that LLMs require localized environmental data to differentiate between similar tropical conditions. * Linguistic shifts pose a significant challenge, as model performance dropped by approximately 10% when processing the French version of the dataset compared to the English version. ## Optimization and Reasoning Strategies * Implementing Chain-of-Thought (CoT) prompting—where the model is directed to explain its reasoning step-by-step—led to a significant 10% increase in diagnostic accuracy. * Researchers utilized an LLM-based "autorater" to scale the evaluation process, scoring answers as correct if the predicted diagnosis was meaningfully similar to the ground truth. * In tests regarding social biases, the study found no statistically significant difference in performance across race or gender identifiers within this specific TRINDs context. * Performance remained stable even when clinical language was swapped for consumer-style descriptions, suggesting the models are robust to variations in how patients describe their symptoms. To improve the utility of LLMs for global health, developers should prioritize the inclusion of regional risk factors and location-specific data in prompts. Utilizing reasoning-heavy strategies like Chain-of-Thought and expanding multilingual training sets are critical steps for bridging the performance gap in underserved regions.

figma2 min readCurated summary

Welcome to The Prompt | Figma Blog

AI may transform design and building, but its ultimate impact remains unsettled. Figma’s *The Prompt* explores that uncertainty through essays and interviews with experts across design, engineering, product development, and the built environment. The collection argues that human judgment—especially the ability to ask thoughtful, well-framed questions—will remain central to making AI useful. ## Prompting as a Creative Discipline - Prompt engineering is described as the practice of getting better answers by asking better questions. - Like interviewing or editing a magazine, effective prompting requires: - Clear context - Thoughtful framing - Useful guidance - AI’s capabilities are treated as largely inert without human direction; people must coax useful results from the technology. - The act of questioning is presented as a fundamentally human and creative instinct. ## The Purpose of *The Prompt* - Created by Figma’s Story Studio and Brand Studio, the magazine launched at Config 2024. - It combines writing, interviews, and illustrations to examine how AI is changing creative and technical work. - Contributors come from both inside and outside Figma and work across: - Design - Engineering - Product development - Robotics - Manufacturing - Residential housing ## Questions About AI’s Future The magazine uses a range of prompts to investigate both immediate applications and larger societal questions, including: - What constitutes good design when AI can generate and automate more work? - Whether code becoming a commodity should be feared - How much data is actually necessary - How starting with imperfect or incomplete ideas can shape innovation - Whether AI development can move beyond technological echo chambers - If efficiency undermines creativity - How to build AI features that people both want and trust - The relationship between artificial design intelligence (ADI) and artificial general intelligence (AGI) - Whether automation can unlock the full potential of design systems - The role of robots in construction and housing - Whether society is entering an age of androids ## Practical and Long-Term Perspectives - Contributors examine ambitious challenges, such as applying AI to manufacturing and housing. - They also focus on what AI can deliver reliably today rather than only speculating about distant possibilities. - The goal is to make complex systems more understandable and usable while learning how to guide AI more effectively. Figma presents *The Prompt* as both a magazine and an experiment in inquiry: meaningful progress with AI depends not just on increasingly capable systems, but on humans asking clearer, more imaginative, and more responsible questions.

Read original(opens in new tab)
figma2 min readCurated summary

David Hoang on how AI brings design and development together | Figma Blog

AI is reshaping creative tools by making interfaces more dynamic, multimodal, and less fully controlled by designers. David Hoang argues that design and engineering are converging as AI augments both disciplines, enabling more people to move from learning and prototyping to shipping real products. Replit’s focus is “Artificial Developer Intelligence” (ADI), aimed at increasing autonomy, productivity, and collaboration rather than pursuing general intelligence. ## Dynamic Interfaces and a New Design Paradigm - Like the rise of mobile, AI is resetting the field and giving designers an opportunity to rethink how people interact with technology. - Interfaces will increasingly adapt across modalities and form factors instead of presenting a fixed, fully designed experience. - Designers must relinquish some control over presentation while shaping how systems behave and respond. - AI, spatial computing—including Apple Vision Pro—and other emerging technologies are converging to create new interaction models. ## Design and Engineering Converge - AI can augment both designers and engineers, making the boundaries between the disciplines less distinct. - Product development is evolving toward a tightly integrated design-and-engineering practice. - Replit frames this strategy as Artificial Developer Intelligence, or ADI, rather than Artificial General Intelligence. - ADI is intended to give people greater autonomy and productivity through collaboration between humans and AI. ## Replit’s Vision for Artificial Developer Intelligence - Future ADI agents could: - Generate code and complete code automatically. - Build complex software architectures. - Orchestrate advanced tools deployed on Replit. - Understand how teams work and improve organizational collaboration. - Hoang describes Replit as a potential “technical co-founder” for people with ideas but limited technical expertise. - The goal is to accelerate the path from learning to coding, launching a business, and scaling an idea. ## Lowering the Barrier to Building Software - AI-powered tools can enable nontechnical users to create technically sophisticated applications. - A Replit hackathon example showed a nontechnical product manager producing work that surpassed projects built by teams of engineers. - Prompting becomes an important skill because translating goals into effective instructions resembles the product manager’s role in defining what should be built. AI’s practical impact may be greatest when it helps people combine product judgment, design, and engineering execution. Rather than replacing creative or technical roles, tools like ADI are positioned to expand who can build software and shorten the distance between an idea and a working product.

Read original(opens in new tab)
figma2 min readCurated summary

Shipping Hype: PMs on What it Takes to Bring AI Features to Market | Figma Blog

AI’s rapid rise has pressured companies to launch features quickly, but hype alone does not produce useful products. Product leaders at Figma, Asana, Duolingo, and LinkedIn argue that successful AI development starts with real user problems, clear definitions, and realistic expectations about current models. AI should be treated as a tool for improving valuable workflows—not as a solution looking for a problem. ## Start with User Problems - Teams should identify user needs before deciding whether AI belongs in the feature. - Figma PM Conor Woods recommends asking: - Can the problem benefit from a large existing data set? - Is some margin of error acceptable? - Is AI genuinely improving the experience, or merely hiding poor UX? - LLMs are well suited to tasks such as organizing information and generating summaries, but they are unreliable when perfect accuracy is required or when they must invent entirely new experiences. - AI-generated inaccuracies and hallucinations are unavoidable with current models, making AI inappropriate for high-stakes, precision-critical tasks. - Asana uses a simple test: does the feature save users meaningful time? - Its Smart Status feature drafts project updates, reducing a task from roughly 20 minutes per week to two minutes and making the return on investment immediately clear. ## Specify the Problem Precisely - Generative AI can serve many different underlying needs, which makes vague feature descriptions dangerous. - Saying “we’ll summarize text” leaves open important questions about the user’s actual goal. - A user might want a summary to: - Understand a document’s subject - Identify action items - Extract decisions or other specific information - Product teams need to define the desired outcome and detailed use case rather than relying on broad descriptions of AI capabilities. - Greater specificity helps designers, engineers, and stakeholders develop a shared understanding of what the feature should do. AI features are most effective when they address a concrete, measurable user problem and acknowledge the limits of current models. Teams should define the user outcome first, then determine whether AI is the appropriate and trustworthy way to achieve it.

Read original(opens in new tab)
figma2 min readCurated summary

Introducing AI to FigJam | Figma Blog

FigJam’s new AI features are designed to solve practical collaboration problems rather than serve as a novelty. Users can generate meeting templates and diagrams from plain-language prompts, summarize brainstorms, and automatically organize sticky notes. Figma argues that this lowers the barrier to visual collaboration while helping experienced users move more quickly from ideas to action. ## AI-Powered FigJam Features - Generate templates for weekly syncs, brainstorms, reviews, and other meetings. - Create visual timelines and organizational charts from a simple prompt. - Summarize the contents of a brainstorm or meeting. - Sort and group sticky notes by theme automatically. - Customize generated outputs based on common workflows and best practices. ## Lowering the Barrier to Visual Collaboration - FigJam AI lets users describe their goals in everyday language instead of learning specialized design software. - A prompt such as “I need a meeting with four people” can produce an initial meeting template. - This approach makes visual collaboration more accessible to people without design backgrounds. - It supports Figma’s goal of “lowering the floor and raising the ceiling”: making the product easier to use while expanding what users can accomplish. ## Solving the Blank Canvas Problem - Starting with an empty FigJam file can make users unsure how to begin. - AI acts as an initial brainstorming partner, helping users move toward actionable next steps. - Tasks such as summarizing complex discussions or synthesizing ideas into categories can take significant manual effort. - Automating this work allows teams to focus on discussion, decision-making, and higher-level collaboration. ## Building AI Around Real User Problems - Figma says its product team drew on its own experience using FigJam to identify useful applications. - The features focus on everyday collaboration needs rather than adding AI for its own sake. - Templates and prompts are based on established practices and common use cases. - The article presents generative AI as a way to make visual tools more useful and approachable across disciplines. FigJam AI is positioned as a practical assistant for getting started, organizing information, and reducing repetitive work. Its main value is helping more people participate in visual collaboration without requiring them to master design tools first.

Read original(opens in new tab)