Loading memory…
Loading memory…
Velvet is a data research company building the datasets that power the next generation of multimodal AI. Founded by Lucas Mantovani (ex Meta FAIR) and Lucas Tucker (ex Adobe Infrastructure), our mission is to make AI more human by producing high-quality audiovisual training data for frontier labs. We're hiring a Founding Machine Learning Engineer to build the pipelines that turn raw footage into clean, structured training data. This is a hands-on, execution-heavy role at the intersection of ML engineering and research. You'll own the full lifecycle — from writing and testing processing scripts to deploying them at scale across thousands of hours of video. Build and enhance post-processing pipelines that clean, validate, and package large volumes of video and audio data for multimodal model training. These pipelines must handle wide variation in speech, visual quality, and format — making robustness a huge engineering challenge. Deploy and fine-tune open-source models for speech recognition, speaker diarization, video segmentation, and related tasks. Design infrastructure for large-scale distributed processing — parallelizing thousands of compute jobs across cloud platforms and opti