What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]
<!-- SC_OFF --><div class="md"><p>We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI</p> <ul> <li>Studio quality speech/audio datasets (high fidelity recordings)</li> <li>Egocentric household activity video datasets (first person daily task recordings)</li> </ul> <p>One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.</p> <p>Some of the recurring challe
Story Overview
We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI
- Studio quality speech/audio datasets (high fidelity recordings)
- Egocentric household activity video datasets (first person daily task recordings)
One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.
Some of the recurring challe