ML engineer, nine years in speech and audio.
I work on real-time lip sync — audio-driven talking heads, and the inference work that gets them running in real time: 2-step distillation, distilled VAE, FP8 quantization, TensorRT. Before that, production TTS and STT serving on multi-GPU infrastructure, streaming pipelines with endpointing, voice activity detection and CTC decoding, and on-device ASR and NLU under embedded constraints.
| TFHubert / TFWav2Vec2 | TensorFlow implementations of Meta's self-supervised speech models, in huggingface/transformers |
| vadkit | Multi-provider voice activity detection — FireRedVAD, Silero, FSMN-VAD, WebRTC — behind one streaming API for endpointing |
| pe-av-syncnet | SyncNet on Meta's Perception Encoder audio-visual representations, for lip-sync evaluation |
| whisper-rl | Reinforcement learning on Whisper, past what supervised fine-tuning reaches |
| denoisers | PyTorch waveform denoisers for speech enhancement |
| spokestack-python | Python library for embedded speech: on-device wake word, ASR and natural language understanding |
| ai-agent-security-2026 | Kaggle gold — red-teaming LLM agents for multi-step tool-misuse attacks |
Available for consulting on speech, ASR/TTS and inference optimization.





