Randomized Controlled Trial

2 posts

figma2 min readCurated summary

Measuring Time Savings From Figma Make | Figma Blog

Figma’s Data Science team found that Figma Make reduced design-task completion time by 20% and made work 16% easier. Product managers benefited most, completing tasks 23% faster and reporting a 37% improvement in ease. Because ordinary A/B tests and observational analyses could not adequately control for task complexity and user experience, Figma used a randomized controlled trial (RCT) with 100 participants. ## Why Measuring AI Time Savings Is Difficult - Productivity is influenced by confounders such as: - Job tenure and career experience - Individual design ability - Task complexity - Without controlling for these factors, it is difficult to determine whether improvements come from AI or from differences among users and tasks. ## Limitations of Common Research Methods - **Online A/B testing** - Randomly assigning users to treatment and control groups helps balance user characteristics. - However, users may perform different tasks, making it difficult to ensure that task complexity is comparable. - **Causal inference using product logs** - Methods such as propensity score matching require all relevant confounders to be present in the data. - Anonymized logs cannot capture subjective factors such as a user’s design experience. - Instrumental-variable analysis requires a valid factor that influences AI usage without independently affecting task speed; Figma could not identify one. ## The Randomized Controlled Trial - RCTs were selected because they can control confounders before data collection begins. - The study combined: - Random assignment to Figma Make and control groups - Identical tasks for all participants - Moderation by trained researchers - The study focused only on Figma Make to avoid introducing variables from multiple AI tools. - Participants included 100 people: - 50 product designers - 50 product managers - The sample size was based on effect sizes from prior industry research, including GitHub Copilot RCTs, followed by a statistical power analysis. ## Findings - Overall, Figma Make: - Made design work **20% faster** - Made work **16% easier** - Product managers experienced the largest gains: - Tasks were **23% faster** - Tasks were **37% easier** The study suggests that a carefully controlled RCT is a more reliable way to measure AI’s productivity impact when task differences and user characteristics are difficult to capture in product data. Teams evaluating similar tools should standardize tasks, randomize participants, and moderate the study to separate genuine AI benefits from other sources of variation.

Read original(opens in new tab)
google3 min readCurated summary

Collaborating on a nationwide randomized study of AI in real-world virtual care

Google and Included Health plan to launch, pending IRB approval, a nationwide randomized study of conversational AI in real-world virtual care. Unlike prior simulated or small feasibility studies, it will prospectively evaluate AI with consented patients across varied conditions and locations, comparing it with standard clinical practice. The goal is to generate rigorous evidence about safety, usefulness, limitations, and impact on patients and clinicians. ## Moving from Simulation to Real-World Evaluation - Earlier research demonstrated clinician-level capabilities in simulated consultations and retrospective analyses. - A feasibility study with Beth Israel Deaconess Medical Center began testing conversational AI in clinical workflows, using measures such as safety-supervisor interruptions. - The new study will advance beyond feasibility through: - A randomized controlled design - Nationwide recruitment - Consented participants - Real patients, clinical concerns, and virtual-care workflows - Controlled comparison with standard practice ## A Phased Approach to Medical AI Research - Google argues that medical AI should be evaluated with evidence standards similar to other medical interventions. - Each research phase adds information about: - Patient and clinician experiences - Safety - Usefulness - The AI system’s capabilities and limitations - Results from each stage are intended to guide safer, more responsible development and deployment. ## Foundational Research Behind the Study ### Diagnostic and Management Reasoning - The AMIE system was developed to handle medical interviews and clinical reasoning. - Studies with patient actors and synthetic cases found that AMIE could match or exceed primary care physicians in simulated diagnostic accuracy and conversation quality. - Later work expanded the system to: - Longitudinal disease management - Clinical-guideline and patient-history reasoning - Investigation and treatment planning - Interpretation of multimodal evidence ### Personalized Health Insights - Research on the Personal Health Agent examined how AI could interpret personal health data, including sleep and activity information from wearables. - Its multi-agent architecture combined the roles of: - Data scientist - Medical domain expert - Health coach - This work informed Fitbit Labs tools such as Symptom Checker and Medical Records Navigator and Plan for Care. ### Navigating Health Information - Google’s “wayfinding” AI research explored how conversational agents can help people find and understand health information. - The system uses proactive guidance, goal recognition, and tailored conversations to make health information searches more practical and useful. ## Practical Conclusion The partnership with Included Health represents a transition from demonstrating what medical AI can do in controlled environments to measuring how it performs at scale in actual care. A nationwide randomized trial could provide the evidence needed to determine whether conversational AI can safely improve virtual care and expand access to medical expertise.

Read original(opens in new tab)