OpenWALDO aims to blow the doors off proprietary AI training models
- ID
- 13695
- Status
- summarized
- Published
- 13 Aug 2026, 12:57 AM
- Fetched
- 13 Aug 2026, 6:05 AM
- Provider
- The Register
- Category
- technology
- Original URL
- https://www.theregister.com/ai-and-ml/2026/08/12/openwaldo-aims-to-blow-the-doors-off-proprietary-ai-training-models/5286864
- Source URL
- https://www.theregister.com/headlines.atom
Summary
- Score
- 5.5
- Created
- 13 Aug 2026, 6:06 AM
- Tags
- Audience
- ai_ml_learnersdeveloperssaas_founders
What happened
Gregory Kurtzer, founder of CentOS and Rocky Linux, has launched OpenWALDO—a project to build a shared, open-source AI training dataset with full provenance and a bill of materials. Funded by his AI infrastructure company CIQ, the effort currently targets 167 billion transparent tokens, a fraction of the trillions used by major AI labs. The project argues that even 'open-weight' models hide their training data, creating legal and compliance risk for downstream users.
Why it matters
If you ship products using open-weight models, you currently have no auditable trail for training data lineage—OpenWALDO's 'bill of materials' concept could eventually let you point to a verified baseline corpus and reduce copyright/consent exposure. But at 167B tokens today versus trillions in proprietary datasets, this is not yet something you can train a competitive model on; treat it as a project to watch, not a dataset to use.
Discussion angle
The gap between 167B transparent tokens and the trillions used by frontier labs—is an open training dataset a meaningful compliance tool for builders, or is it too small to matter for anyone shipping real AI products today?