VERIFIED ENGLISH WORKPLACE
Narrator - Videos for Data Collection Project, AI training
Toptal β’ Anywhere in the World (Remote)
100% English First
Est. β¬72,000 - β¬95,000 / year
Headquarters:
About the Client
Our client is involved in an innovative data-collection project, which focuses on creating video datasets for training AI models. These AI models are designed to learn from first-person (POV) video footage. The captured videos depict real people involved in hands-on tasks with a mounted phone on their head or chest, paired with spoken narration. The purpose is to enable the model to connect visual inputs (objects, hands, actions) with audio descriptions.
Sub-areas for the videos
- Home & Daily Tasks: cleaning; laundry (sorting, folding); organizing a closet; house tours; pet care; packing luggage; loading appliances / dishwashing
- Repairs & DIY: home repair; gardening / farming; furniture assembly; woodworking; plumbing; electrical work; bicycle maintenance; soldering electronics
- Textile & Craft Arts: sewing; knitting; crocheting; using a loom; leathercrafting; bookbinding; crafting jewelry; pottery / ceramics
- Professional Trades & Industrial: automotive repair / maintenance; warehousing / logistics; construction / woodworking; laboratory work; operating heavy machinery controls; assembly-line packaging
- Hobbies & Arts: art β drawing; art β ceramics; art β painting; playing instruments; model building; outdoor survival / camping; calligraphy
- Technology & Computing: using computers or devices (e.g. audio mixers); gaming; product demos; VR / AR interaction; 3D printer setup / maintenance
- Outdoors & Activity: city tours; shopping; navigating public transit
- Personal Care: haircut; applying makeup; detailed grooming routines
- Specialized & Professional: medical procedures; first aid training; professional barista workflows; culinary chef / knife work; lab protocols (pipetting, titrations)
About the Role - Narrator
Watches the POV video and describes the scene and task in first person, in English, as if explaining it to a model that can't see. Fluency and describing ability matter most; knowing the activity is not required.
- Deliverable β audio narration synced to the video, recorded after the footage
- Language β English; fluency is criterion #1
- Density β at least 25 words/min on average, no long silences
- Content β 90%+ about the setting or task steps
Required Skills
- English fluency
- Ability to narrate pre-recorded videos
- All details on the Acceptance Criteria will be shared prior to the work.
Rules
- No personal data on video β third-party faces, screens, documents, plates, addresses, mirror reflections.
- 100% human narration β no TTS or synthetic voice.
Engagement Details
- Compensation model: Hourly paid, considering the total number of hours for approved videos.
- Commitment Type: Flexible hours based on video submissions.
- Duration: Ongoing, with earnings based on the number of approved videos.
- Location: Remote, with flexible overlap with client's timezone.
- Start Date: Immediate start upon approval of test video.