experiment-design

3 posts

discord

Measure Less to Learn More: Using Fewer, Higher-quality Metrics to Capture What Matters (opens in new tab)

Discord argues that experiments should measure fewer, higher-quality metrics rather than automatically collecting every potentially useful signal. Large metric sets increase compute and cognitive costs while creating a tradeoff between false positives and missed real effects. Multiple-testing corrections such as Benjamini–Hochberg reduce false discoveries but also lower recall, so the most effective solution is selecting metrics that represent distinct, important concepts. ## The Cost of Measuring Too Much - Discord’s “Default Metric List” gradually expanded as teams added metrics and rarely removed them. - More metrics create: - Higher compute costs - More difficult experiment readouts - Increased risk of false positives - With 100 metrics and an uncorrected significance threshold of 0.05, roughly five metrics may appear significant purely by chance. - Correcting for multiple comparisons reduces false alarms but makes genuine changes harder to detect. ## The Multiple Comparisons Problem - Discord uses the Benjamini–Hochberg (BH) procedure to control the false discovery rate at 5%. - BH ranks p-values and compares each one with a rank-specific threshold: `i × α / n` where `i` is the metric’s rank, `α` is 0.05, and `n` is the total number of metrics. - A metric with an unadjusted p-value of 0.038 might be significant without correction but fail after BH adjustment. - BH treats all metrics equally because it has no information about which ones are more likely to reflect a real effect. - The resulting tradeoff is: - Fewer false alarms - Lower recall for real changes - The article notes that Bayesian methods could potentially incorporate prior knowledge, but Discord’s default system is frequentist. ## Simulation Results - Discord simulated 50,000 experiments containing: - Twenty null metrics generated from `N(0, 1)` - One metric with a real effect centered at `z = 2.8` - The simulations tested how metric count affects: - Experiment-level false alarm rates - Recall of the metric with the known effect - Without correction, false alarm rates rose sharply as more metrics were added—approximately from 23% with five metrics to 93% with 50. - BH kept false alarm rates near 5%, but recall declined as the metric pool grew, falling from roughly 60% to 30% across the same range. - These results demonstrate that adding metrics makes statistical correction stricter and makes genuine effects harder to identify. ## Fewer, Higher-Quality Metrics - Reducing the metrics automatically included in experiments improves the balance between false alarms and recall. - Metrics should be selected for quality and conceptual distinctness rather than added “just to be safe.” - The article’s central conclusion is that no sophisticated statistical method eliminates the underlying tradeoff created by excessive measurement. Teams should maintain a focused default metric set, regularly remove low-value or redundant metrics, and reserve specialized metrics for experiments where they are genuinely relevant.

toss

From Intern to Solo Designer: Growth (opens in new tab)

As a Toss Bank product design intern, Jeon Nuri designed experiments to improve non-member sign-up conversion. She prioritized the funnel using speed and impact, studied previous experiments, and learned that clear, narrowly defined hypotheses were more valuable than constantly generating new ideas. The experience showed that failed experiments can still guide better decisions when they produce actionable learning. ## Prioritizing the Right Funnel Stage - The largest drop-offs occurred in the intro, consent, and identity-verification screens. - Consent and identity verification were shared modules requiring legal and compliance review, making rapid iteration difficult. - The intro screen could be changed more quickly and had the potential to affect the greatest number of users. - Based on this speed-versus-impact assessment, she chose the intro screen as the starting point. ## Learning from Previous Experiments - Instead of immediately designing new concepts, she reviewed existing experiments, including both winners and unsuccessful variations. - She examined: - The problem each experiment addressed - The reasoning behind its hypothesis - How the test variation was designed - Experiments from unrelated screens were also useful because their problem definitions and hypothesis structures could be adapted. - The main lesson was that inexperienced experimenters benefit more from systematically analyzing existing learning than from rushing to create new ideas. ## First Experiment: A Counselor Concept - The first variation presented benefits as if they were being recommended by a counselor and offered a small number of choices. - The hypothesis was vague: fewer choices would increase conversion. - The result was negative: - Click-through rate fell by more than 10%. - Conversion rate fell by more than 3%. - The design actually introduced more choices than the original, which had only one CTA button. - The experiment also failed to consider why users had entered the screen and whether they needed recommendations. - This led her to analyze the existing screen and user context before creating a hypothesis. ## Identifying and Solving Concrete Problems - Rather than inventing an entirely new design, she identified two specific weaknesses in the existing version: - The copy did not clearly communicate benefits users cared about. - Images loaded slowly, taking two to three seconds on low-end devices. - Previous experiments showed that users responded well to messages about high interest rates and receiving interest daily. - She incorporated those themes into the copy and optimized the visuals with newer graphics and lower-weight image formats. - Both click-through rate and conversion rate increased, demonstrating that a hypothesis grounded in clear problems can provide a stable direction for design. ## Making Benefits Easier to Imagine - Building on the earlier results, she changed functional wording into language that helped users imagine a concrete situation and immediate benefit. - Instead of simply explaining that interest could be earned after depositing money for one day, the revised copy foregrounded the moment when users would experience the benefit. - Copy alone increased CTR by 5% and also produced a meaningful improvement in CVR. - The result reinforced that different expressions of the same information can create significantly different first impressions. ## Principles for Designing Experiments - Break the funnel into stages and prioritize opportunities by speed and potential impact. - Understand the existing context before defining the core problem. - Study previous experiments through their hypotheses and problem definitions, not just their numerical outcomes. - Establish a clear hypothesis and success metric before designing the variation. - Make sure the experiment visibly tests the stated hypothesis. - Treat failure as input for the next decision rather than as wasted effort. A practical starting point for new designers is to begin with a small, focused experiment—but make the hypothesis precise enough to guide both the design and the next iteration.

stripe

Businesses grow revenue on Stripe 27 percentage points faster after accepting financing through Stripe Capital (opens in new tab)

Stripe’s two-year randomized trials found that businesses accepting Stripe Capital financing grew faster than comparable businesses without financing. The 2023–2025 study showed an average 27-percentage-point growth advantage, while the fastest-improving 10% saw an average boost of 211 percentage points. The results suggest that embedded, data-driven financing can help small businesses overcome traditional lending barriers and invest in growth. ## Proving Financing Causes Growth - Stripe compared businesses that accepted Capital with similar businesses matched on credit, revenue, and longevity. - The study was conducted across two periods: - **2020–2021:** financing was associated with a 114-percentage-point average growth boost, though pandemic-era economic conditions may have influenced results. - **2023–2025:** financing still produced a strong 27-percentage-point average boost in a different economic environment. - Stripe conducted the trial at scale, serving 76,000 financed businesses in 2025 alone. ## Strongest Effects Among Small Businesses - Businesses processing **$3,000–$76,000 annually** saw average growth-rate improvements of **33–43 percentage points**. - Businesses processing less than **$52,000 annually** with top-tier credit scores saw even larger improvements of **94–106 percentage points**. - Even businesses with low or unavailable credit scores experienced **11–18 percentage-point** growth improvements. - Stripe says its data-driven process delivers financing in **1–2 days**, compared with roughly **14–40 days** at traditional banks. - Traditional bank applications are often time-consuming, and rejection rates can approach 50%, including for established businesses. ## Growth-Oriented Spending Produces Better Results - A survey of approximately 900 participating businesses found that financing use strongly correlated with outcomes. - Among businesses with top-tier credit, those using funds to launch products, start projects, or scale operations saw average growth boosts of **70–95 percentage points**. - Examples included: - MyPark used financing to deploy additional revenue-generating machines. - Xirsys expanded server infrastructure into China, India, and Japan, more than doubling annual revenue. ## Expanding Access Through Embedded Finance - The World Bank estimates a **$5.7 trillion** funding gap for SMBs in developing economies. - Platforms that already manage payments or business operations can use transaction data to make proactive financing offers. - This model broadens access beyond traditional credit scoring and may encourage owners to pursue investments they would otherwise avoid. - Marketplaces and software platforms are positioned to become important channels for closing the global SMB funding gap. Stripe’s research supports using embedded, data-based financing to provide faster access to capital, particularly for small businesses and owners pursuing concrete expansion plans. However, financing remains subject to approval and may take the form of loans or merchant cash advances depending on the market.