Computer Vision Engineer
Industrial Video Analytics · Object Detection · Object Tracking · Pose, PPE & Zone Analytics · OCR
I build multi-camera vision systems from dataset design and held-out evaluation through RTSP integration, tracking, pose logic, field iteration, and event evidence
LinkedIn · Email · Resume · IEEE Xplore · Interactive Matrix profile
Crane operator safety vision system
Owned a six-camera system end to end: built a 1,163-image / 7,499-object dataset, trained YOLOv8m, and reached 0.932 precision, 0.878 recall, 0.920 mAP@0.50, and 0.707 mAP@0.50:0.95 on a held-out test set. Implemented active-view selection, ByteTrack identities, pose-based wrist-to-pendant association, helmet checks, temporal voting, and saved event evidence; the work became a paid client customization.
Industrial safety video analytics
Co-developed modular RTSP analytics for PPE, danger zones, tracking and counting, equipment use, and worker activity. Offline PPE validation reached 0.906 mAP@0.50 and 0.772 mAP@0.50:0.95 on 579 images / 1,748 objects.
Multi-camera pilots and low-resolution CCTV
Reviewed 17,696 CCTV frames from eight camera feeds and curated 1,156 relevant frames for a multi-camera industrial pilot, with view-specific zones designed to reduce long-range and duplicate detections. Built a worker-activity prototype using detection, tracking, pose, fixed zones, temporal stabilization, repeat-alert suppression, reason labels, and image/video evidence.
Image preprocessing at Shtrih-M
Developed and visually evaluated a Python/OpenCV workflow across 1,147 raw GRBG Bayer images captured under varied lighting, glare, backgrounds, and packaging conditions. Documented seven processing functions, before/after comparisons, RGB histograms, adaptive decision logic, and implementation details in a 64-page technical report.
NeuroQuest — Automatic generation of neurocomics
Built an end-to-end RTSP-to-comic system combining detection, tracking, face and action recognition, scene interpretation, language and image generation, and PDF assembly. First author, lead writer, and oral presenter at ACDSA 2024.
video → perception → scene understanding → narrative → generated comic
Jelluvi — a local-first browser extension for exporting AI conversations to nine portable formats without telemetry, accounts, or remote rendering, demonstrating TypeScript, browser APIs, testing, and privacy-conscious product engineering.
- Vision: object detection, object tracking, multi-camera analytics, pose estimation, PPE and zone analytics, OCR, image preprocessing
- ML: Python, PyTorch, OpenCV, Ultralytics YOLO, TensorFlow, ByteTrack, Deep SORT, NumPy
- Systems: RTSP, FastAPI, Docker, PostgreSQL, GPU inference, dataset design, held-out evaluation, error analysis
Educational and historical public archive
- Word Solver CV — personal educational Computer Vision/OCR portfolio artifact
- Specific questions of data analysis
- Data science competitions
- Applied information security tasks
- Technical article: computer vision word solver
These repositories preserve earlier learning and public engineering history.
Vlad Voropaev · Computer Vision Engineer · UAE

