Research Scientist / Engineer – Training Infrastructure at Luma AILuma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. We are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change. We are looking for engineers with significant experience solving hard problems in PyTorch, CUDA and distributed systems. You will work alongside the research team to build and train cutting‑edge foundation models on thousands of GPUs that are designed to scale from the ground up.Responsibilities Design, implement, and optimize efficient distributed training systems for models with thousands of GPUsResearch and implement advanced parallelization techniques (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel)Build monitoring, visualization, and debugging tools for large‑scale training runsOptimize training stability, convergence, and resource utilization across massive clustersExperience Extensive experience with distributed PyTorch training and parallelism in foundation model trainingDeep understanding of GPU clusters, networking, and storage systemsFamiliarity with communication libraries (NCCL, MPI) and distributed system optimization(Preferred) Strong Linux systems administration and scripting capabilities(Preferred) Experience managing training runs across >100 GPUs(Preferred) Experience with containerization, orchestration, and cloud infrastructureSeniority Level Mid‑Senior levelEmployment Type Full‑timeJob Function OtherIndustries Software Development#J-18808-Ljbffr