Language Models (VLMs), Multimodal AI, and production-grade AI deployment. The role will focus on designing, developing, optimizing, and deploying advanced vision and multimodal AI solutions for enterprise-scale applications.
The ideal candidate should have deep expertise across object detection, image segmentation, image classification, OCR, image captioning, scene understanding, pose estimation, depth estimation, facial recognition, and camera-based inference pipelines. Strong experience with modern computer vision architectures such as YOLO, Faster R-CNN, SSD, U-Net, Mask R-CNN, DeepLab, ResNet, EfficientNet, and Vision Transformers is expected.
No data