
NVIDIA opens Nemotron agent-training data: 10T+ tokens and an interactive prompt atlas
NVIDIA released its Nemotron open-data collection on Hugging Face — over 10 trillion pre-training tokens plus millions of post-training samples, an interactive Prompt Atlas that maps each sample by domain and tool use, and Nemotron-Personas now spanning 10 countries and 2.4B people. The signal is inspectability: open datasets, not just open weights, let developers trace why an agent calls a tool or recovers from a failure. It also deepens NVIDIA's play to seed its GPU ecosystem — nearly 145 ICML papers already cite Nemotron models and data.
Source: huggingface.co ↗