Curated summary
Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph
Netflix’s growing use of machine learning across personalization, Studio, payments, advertising, and other domains has created a fragmented ecosystem of tools and metadata. The Metadata Service (MDS) addresses this problem by building a Model Lifecycle Graph that connects models, features, pipelines, experiments, datasets, and ownership information. Its goal is to make ML assets discoverable, understandable, and reusable across organizational boundaries.
A Fragmented Machine Learning Landscape
- Netflix ML has expanded from personalization into areas such as:
- Studio production and post-production
- Fraud detection and payment optimization
- Advertising and real-time targeting
- Each domain uses different technologies, metrics, and organizational structures.
- Valuable assets often remain isolated in specialized systems.
- For example, Studio-generated content embeddings could support:
- Contextual ad matching
- Episodic merchandising
- Recommendations based on tone, topic, or mood
- Practitioners struggle to answer basic questions because relevant information is split across:
- Model registries
- Pipeline orchestrators
- Experimentation platforms
- Feature stores
- Dataset systems
- This fragmentation makes discovery, lineage tracking, impact analysis, and ownership difficult.
The Challenge of Connecting ML Infrastructure
- MDS must unify metadata from many independent systems, including:
- Pipeline execution and transformation data
- Model versions, artifacts, deployments, and staleness
- A/B test configurations
- Feature definitions and usage
- Dataset creation and discovery
- User, team, and organization information
- These systems use different identifiers, formats, and conceptual models.
- The core challenge is transforming heterogeneous metadata into a common entity model and connected graph—not merely creating a consolidated user interface.
The Model Lifecycle Graph
- Netflix’s Metadata Service indexes ML-related assets and materializes relationships between them.
- It supports real-time metadata ingestion and cross-domain questions such as:
- Which experiments use a particular model?
- Which models depend on a feature?
- What data sources feed a model?
- Who owns each part of the workflow?
- The graph is intended to make every ML asset discoverable and reusable regardless of its originating team or business domain.
Core Concepts and Vocabulary
- Component: Any uniquely addressable object identified by an AIP URI, such as:
aip://model/registry/ranking-v5aip://user/identity/aliceaip://pipeline/orchestrator/weekly-training
- Entity: A component enriched with properties such as name, description, creation date, and ownership.
- Entity type: A group of entities sharing the same data shape and required properties.
- Domain: An abstract interface for a category of ML assets, such as Models or Pipelines.
- Provider: A concrete backend implementation of a domain, such as Netflix’s internal model registry.
- Separating domains from providers allows multiple systems to implement the same interface without changing how consumers interact with MDS.
- URI-based addressing gives services a consistent way to reference assets and resolve them to connected metadata.
From Events to a Queryable Graph
- MDS receives metadata events through Kafka and AWS SNS/SQS.
- Source systems emit lightweight events containing an event type and resource identifier.
- For example, a model registry might emit a
model_instance_createdevent with the new instance’s ID. - This keeps event producers simple while allowing MDS to enrich events, construct entities, and infer relationships such as connections between models and A/B tests.
The Model Lifecycle Graph provides Netflix with a common layer for connecting previously isolated ML systems. By standardizing identifiers, entities, domains, and providers, MDS can support cross-domain discovery, lineage, impact analysis, and collaboration at scale.
Related reading
Continue with another curated summary.
Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned
Read originalBridging the Gap: Diagnosing Online–Offline Discrepancy in Pinterest’s L1 Conversion Models
Read originalInside the feature store powering real-time AI in Dropbox Dash
Read originalPredicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning
Read original