English summary for screening — check the original posting before applying.
Turing is developing fully autonomous driving AI and is rapidly expanding its GPU compute resources. This role focuses on operating and improving the GPU clusters supporting machine learning, primarily Slurm environments on AWS/GCP, to maximize computational resource utilization.
Must-haves
- Experience operating job schedulers like Slurm
- Experience operating Kubernetes or cloud infrastructure (AWS/GCP)
- Experience with infrastructure automation (Terraform/Ansible)
- Experience operating Linux servers
- Experience with automation using scripting/programming (Python/Go)
Nice-to-haves
- Experience operating and tuning GPU clusters
- Understanding of ML workloads (training and inference)
- Experience with job scheduling and resource optimization
- Experience integrating and operating Slurm x Kubernetes
- Knowledge of distributed training (PyTorch DDP)
- Knowledge of parallel file systems and high-speed networks
- Experience with cost optimization (FinOps)
Tech stack
SlurmAWSGCPKubernetesTerraformAnsiblePythonGo
Work style
Onsite in Tokyo (Heiwajima Office) with a flextime system (core hours 11:00-15:00).
Other notes
Salary ranges from ¥10,000,000 to ¥20,000,000+ annually, depending on experience and role (Senior/Principal Engineer). Salary information is not to be included in initial application documents.