Description
We’re looking for an experienced ML engineer to fine-tune an open-source speech-to-text model for improved transcription accuracy on English and Hindi-English (Hinglish) conversational audio. Responsibilities: * Prepare and clean speech datasets * Fine-tune the ASR model * Evaluate transcription quality (WER/CER) * Optimize inference and document the training pipeline Requirements: * Experience with speech recognition (Whisper, Cohere Transcribe, NeMo, wav2vec2, etc.) * Strong PyTorch/Hugging Face experience * Experience with multilingual or code-switched speech is preferred * Familiarity with WER evaluation and dataset preparation Please include: * Relevant ASR projects you’ve worked on * Models you’ve fine-tuned * Links to GitHub or publications (if available) * Your expected timeline for an initial fine-tuned model