An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Mercor is seeking USA-based voice actors for a single ~4-hour recording session to train AI speech models. Open to newcomers with a clear, expressive voice and a good home/near-professional studio. Voice may be cloned for the CX AI Agent, with recordings used only for this client.
Applicants should have a professional or near-professional setup and be able to deliver multiple takes with varied emotion and tone. This engagement is confidential and time-bound.
Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking USA-based voice actors with native American English accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is open to both experienced voice professionals and newcomers with a naturally clear, expressive voice and a good recording setup — prior voice-acting experience is a plus but not required.
The engagement itself is a single recording session of approximately 4 hours. A second session of similar length may follow depending on the client's needs, though this isn't guaranteed.
Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose.
Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.)
Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation
Maintain consistency in voice, accent, and delivery across recording sessions
Follow detailed recording guidelines (environment, microphone setup, file formatting)
Perform multiple takes with variation in emotion, emphasis, and style when required
Native American English speaker currently based in the United States
A clear, natural, and expressive voice (prior voice acting, narration, or broadcast experience is a plus, not required)
Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.)
Strong command of intonation, diction, and emotional range
Ability to follow scripts precisely while maintaining natural delivery
Availability for a single ~4-hour recording session, with the possibility of a further session of similar length later on
[IMP]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case.
Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets
Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper)
Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm)