How It Works
Data Sources
Connect & Ingest
OCR + source metadata
Tag
AI metadata extraction
Slice
curated datasets
Deliver

Maintained over time: Workflows re-ingest, re-classify, and re-export on a schedule
1
Connect
Point a Data Connector at your document store. No data migration needed.
2
Tag
Define what you want to know about your documents with Tags and Taxonomies, or let AI suggest them. Classification generates Metadata with values, evidence, and confidence for every document.
3
Slice
Build Data Slices that capture the exact subset each use case needs.
4
Deliver
Export slices to Destinations and enrich documents at the source.
5
Maintain
Schedule Workflows so datasets stay fresh as documents change.
Start Building
Quickstart
Go from zero to extracted metadata in under 5 minutes with the Python SDK.
Use Cases
What teams build: RAG context, compliance, document management, governance.
Cookbooks
End-to-end recipes for RAG pipelines, PII detection, and data quality.
API Reference
Every endpoint, generated from the live OpenAPI specification.
Why Teams Use It
Better AI Results
Curate before you compute. Retrieval works on high-quality, relevant knowledge instead of raw document dumps.
Protect Sensitive Data
Detect PII, PHI, and PCI automatically at scale and route sensitive content before it reaches downstream systems.
Reliable Over Time
Scheduled workflows keep datasets fresh, so answers stay grounded in current documents.
