The Hard Layer of AI Data
- Nexus Data Strategy

- Jun 25
- 5 min read
Supply is rushing in. The part that matters is still thin.

Outside Moscone during Databricks Data + AI Summit, San Francisco, June 2026. The robots really are already here.
I just got back from ten days in San Francisco, using Databricks Data + AI Summit as a base and working outward from there: robotics teams, frontier labs, voice companies, data infrastructure players, and the people building the data layer underneath all of it.
A day down in Mountain View. A lot of side events. A lot of conversations.
One thread ran through almost every discussion.
The data market is not short of supply. New supply is rushing in from every direction, but most of it is the commodity layer: synthetic and crowd-collected video in physical AI, the major languages and clean read speech in voice, the public web crawl for LLMs. That layer is abundant, cheap, and shared by everyone, so it no longer differentiates anyone.
What is now scarce is the opposite: data with verifiable provenance and rights, captured in the real deployment environment, and tied to the specific cases where models still fail. Scale at the easy layer is solved. Precision at the hard layer is not.
In physical AI, that shift is becoming visible fast.
Both layers matter. High-volume egocentric data is useful for pre-training and broad capability. It gives models exposure to human action, object interaction, and environmental variation at a scale that would be impossible to collect only through robots.
But that part of the stack is also where supply is rushing in fastest. Cheap, commoditized egocentric data, much of it from South Asia, is likely to make the broad pre-training layer far less scarce over the next 6 to 12 months.
The harder layer is mid- and post-training data: aligned human-robot data, task-specific demonstrations, recovery cases, sensor-rich environments, and data matched to the actual homes, workplaces, warehouses, kitchens, stores, and streets where systems are expected to operate.
A home or workplace in South Asia can absolutely be useful for broad pre-training. But it is not automatically a substitute for the deployment environment. The layouts, objects, materials, appliances, signage, lighting, density, and ways space is used can all differ.
My view is that the pre-training data layer will commoditize quickly. The mid- and post-training layer will not.
That is where the bottleneck is moving.

With Leo 磊 Su at RealMan Robotics, Mountain View, June 2026.
Voice was the surprise of the trip, and the clearest illustration of the same bottleneck shift.
What struck me first was simply how good the leading models already are. I spent time trying to break one of them, Deepgram, with my Irish accent, fully expecting it to stumble. It held up almost perfectly.
That was not a model breakthrough. It was a training-input decision: the model had been trained on native Irish-accented English, so it could handle the accent properly.
That is where voice is going. It is a quality story, not a volume story.
Not one English. Every English.
The supply side can produce a lot of audio. But as base speech models improve, the bottleneck shifts from generic transcription data to native, contextual, domain-specific speech that fixes the cases where models still fail.
The hard version is not what can be scraped, synthetically generated, or cheaply crowdsourced at scale. It is multilingual conversational audio that is actually native.
Real Hindi and its regional variants, for example. Not Americanized Hindi. Not diaspora-accented Hindi. Not clean read scripts. Real conversation, captured in context, with the metadata and rights that make it usable: human-QA'd transcripts, diarization, and clean ASR and TTS permissions.
Then apply the same standard across every other major language.
Then across the long tail of lower-resource languages, where the data exists but is expensive, uneven, and operationally difficult to collect.
And then comes the hardest category of all: domain-specific audio from real professional settings. The kind of speech that only happens behind closed doors.
A doctor, a patient, and an interpreter. Three languages in one exchange. Messy, high-stakes, context-rich, and commercially valuable.
That data is not sitting on the open web. And it does not get easier to find.
Which is the common denominator across physical AI, voice, and frontier model development. The text frontier runs the same way: the public web crawl is abundant and shared by everyone, while the scarce material is licensed, expert, and proprietary corpora, and high-quality human reasoning and preference data.
The most serious teams I met are going deeper into data that simply is not publicly available. Data that has to be accurate, current, robust, legally usable, and privately sourced.
The easy layer is expanding.
The hard layer is still thin.
One final observation.
In every one of these spaces, the companies springing up are making some version of the same claim: better data, better models, faster.
But better how?
For physical AI, does the data reduce the sim-to-real gap? Does it improve task completion outside controlled environments? Does it expose the model to the objects, layouts, lighting, materials, occlusions, and edge cases it will actually face? Does it help a robot recover when the plan fails?
For voice, does the data improve accent and dialect robustness? Does it handle code-switching, overlapping speakers, background noise, domain vocabulary, and real conversational speech rather than clean read scripts? Does it reduce errors in the languages and settings the product actually needs to serve?
That gap, between claim and proof, will decide who survives.
And the proof will not come from a pitch deck. It will show up in the data itself: where it came from, how it was collected, what rights attach to it, how representative it is, how well it maps to the target use case, and whether it actually improves the failure mode the buyer cares about.
That is the real evaluation layer.
The bottleneck is not more data.
It is the right data: hard to source, accurate, legally usable, fit for purpose, and tied to the actual failure modes teams need to solve.
That is where Nexus Data Strategy focuses.
If you are building in physical AI, voice, or frontier models and your bottleneck is hard-to-source data rather than raw volume, I would like to talk.
Same if you are sitting on proprietary operational, behavioral, sensor, audio, or workflow data that could be valuable to AI teams but has not yet been commercialized.
Nexus helps AI teams source real enterprise data for specific AI use cases, faster and without legal or sourcing dead ends. We work with physical AI companies, voice teams, and foundation model labs on the buy side, and with enterprises sitting on proprietary data on the supply side. Start with a free feasibility screen. #AIData #TrainingData #VoiceAI #PhysicalAI #FoundationModels #DataStrategy #EnterpriseData



Comments