English summary for screening — check the original posting before applying.
Third Intelligence is seeking a Research Engineer to join their AI R&D team. This role focuses on developing the vision capabilities essential for their 'Ubiquitous AGI' initiative, aiming to create AI that understands and responds to humans in real-time.
Must-haves
- Experience training image-based foundation models (scratch, continual learning, fine-tuning)
- End-to-End AI system construction experience in image recognition/video understanding
- Deep knowledge of Python, PyTorch, and Distributed Training Frameworks (DeepSpeed, FSDP)
- Deep expertise in image recognition, object detection, video understanding, or Vision-Language Models (VLM)
- Ability to read, implement, and verify the latest image, video, and multimodal papers independently
Nice-to-haves
- Track record of developing and operating frontier models
- Experience in technical discussions in English and global development environments
- Activity in international technical communities
- Expertise in computational graph optimization (C++/CUDA) or AI accelerators (ASICs)
- Development experience with architectures integrating LLMs and visual models (VLM)
- Experience optimizing inference engines for low-latency, real-time performance
Tech stack
PythonPyTorchDeepSpeedFSDPC++CUDA
Work style
Full remote from anywhere in Japan is possible. Mondays and Fridays are recommended in-office days, and monthly AllHands meetings require attendance at the office in Tokyo.
Other notes
Salary is negotiable. Stock option system available.