AI Weekly Malaysia

Back to items Summaries

OpenWALDO aims to blow the doors off proprietary AI training models

ID
13695
Status
summarized
Published
13 Aug 2026, 12:57 AM
Fetched
13 Aug 2026, 6:05 AM
Provider
The Register
Category
technology
Original URL
https://www.theregister.com/ai-and-ml/2026/08/12/openwaldo-aims-to-blow-the-doors-off-proprietary-ai-training-models/5286864
Source URL
https://www.theregister.com/headlines.atom

Summary

Score
5.5
Created
13 Aug 2026, 6:06 AM
Tags
Audience
ai_ml_learnersdeveloperssaas_founders

What happened

Gregory Kurtzer, founder of CentOS and Rocky Linux, has launched OpenWALDO—a project to build a shared, open-source AI training dataset with full provenance and a bill of materials. Funded by his AI infrastructure company CIQ, the effort currently targets 167 billion transparent tokens, a fraction of the trillions used by major AI labs. The project argues that even 'open-weight' models hide their training data, creating legal and compliance risk for downstream users.

Why it matters

If you ship products using open-weight models, you currently have no auditable trail for training data lineage—OpenWALDO's 'bill of materials' concept could eventually let you point to a verified baseline corpus and reduce copyright/consent exposure. But at 167B tokens today versus trillions in proprietary datasets, this is not yet something you can train a competitive model on; treat it as a project to watch, not a dataset to use.

Discussion angle

The gap between 167B transparent tokens and the trillions used by frontier labs—is an open training dataset a meaningful compliance tool for builders, or is it too small to matter for anyone shipping real AI products today?

Top