About the role
structured by ORIBuild and run secure AI research platforms with scalable cloud and GPU compute to accelerate experimentation. Are you looking for an exciting opportunity to join a dynamic and growing team in a fast-paced and challenging area?
What you will do
- Own the architecture, technical roadmap, and day-to-day engineering of the team’s AI research platform.
- Design and operate secure, scalable cloud and CPU/GPU infrastructure, including clusters, workload scheduling, storage, networking, capacity planning, performance, and cost.
- Build reproducible, self-service research environments and platform automation using standardized container images, GPU software stacks, dependency and artifact management, infrastructure-as-code, CI/CD, and orchestration.
- Partner with researchers to translate requirements from LLM, agentic AI, model evaluation, retrieval, inference, and accelerator benchmarking projects into dependable platform capabilities, and prepare research prototypes for downstream integration through packaging, service interfaces, and operatio
- Assess and onboard new AI infrastructure technologies, cloud services, accelerators, and developer tooling through practical benchmarks, architecture reviews, and controlled pilots.
What they are looking for
- Master's degree in computer science, software engineering, computer engineering, information systems, or a related technical field, plus at least 4 years of relevant industry experience building or operating cloud, platform, or infrastructure systems.
- Hands-on experience operating production-grade environments, including EC2, networking, storage, monitoring, and cost management.
- Strong knowledge of Linux systems, containers, networking, storage, and distributed systems, with experience in orchestration or workload-scheduling technologies.
- Experience with infrastructure-as-code and delivery automation using tools such as Terraform or CloudFormation, configuration management, and CI/CD pipelines.
- Proficiency in Python and shell scripting, with sound software engineering practices for building maintainable automation, services, and platform integrations.
- Working knowledge of AI/ML development workflows and demonstrated ability to translate research requirements into platform designs, communicate technical trade-offs, manage delivery dependencies, and drive risk and control items to closure.
- Experience in one or more of the following domains: cloud and accelerated compute platforms (e.g., AWS, CPU/GPU infrastructure, autoscaling, workload scheduling, resource quotas, capacity planning); research environments and developer platforms (e.g., containers, Jupyter, package and dependency mana
Nice to have
- Experience supporting AI/ML research or development teams and operating platforms for GPU-intensive, LLM, agentic, or other distributed AI workloads.
- Experience with Kubernetes or managed container platforms, GPU software stacks, cluster schedulers or distributed-compute frameworks such as Slurm or Ray, observability platforms, and cloud cost optimization.
- Experience delivering infrastructure in a regulated enterprise and working with cybersecurity, architecture, technology risk, and audit stakeholders.
- Familiarity with finance or financial use cases.
Full posting text
Build and run secure AI research platforms with scalable cloud and GPU compute to accelerate experimentation.
Are you looking for an exciting opportunity to join a dynamic and growing team in a fast-paced and challenging area? This is a unique opportunity for you to work with the Global Technology Applied Research (GTAR) center at JPMorgan Chase & Co. The goal of GTAR is to design and conduct research across multiple frontier technologies, in order to enable novel discoveries and inventions, and to inform and develop next-generation solutions for the firm's clients and businesses. As an AI Platform Engineer – Vice President at JPMorgan Chase within the Global Technology Applied Research (GTAR) center, you will engineer, operate, and evolve the shared platforms that enable GTAR researchers to develop, evaluate, and mature AI solutions in secure, scalable, firm-approved environments. You will own the cloud and compute foundation for AI research workloads, build reusable platform services and automation, and partner with researchers and enterprise technology teams to make experimentation reproducible, reliable, well-controlled, and ready for downstream integration. Job responsibilities Own the architecture, technical roadmap, and day-to-day engineering of the team’s AI research platform. Design and operate secure, scalable cloud and CPU/GPU infrastructure, including clusters, workload scheduling, storage, networking, capacity planning, performance, and cost. Build reproducible, self-service research environments and platform automation using standardized container images, GPU software stacks, dependency and artifact management, infrastructure-as-code, CI/CD, and orchestration. Partner with researchers to translate requirements from LLM, agentic AI, model evaluation, retrieval, inference, and accelerator benchmarking projects into dependable platform capabilities, and prepare research prototypes for downstream integration through packaging, service interfaces, and operational-readiness reviews. Assess and onboard new AI infrastructure technologies, cloud services, accelerators, and developer tooling through practical benchmarks, architecture reviews, and controlled pilots. Maintain reliable, well-controlled platforms through observability, incident response, patching and lifecycle management, access and data controls, risk and control evidence, operational documentation, and runbooks. Required qualifications, capabilities, and skills Master's degree in computer science, software engineering, computer engineering, information systems, or a related technical field, plus at least 4 years of relevant industry experience building or operating cloud, platform, or infrastructure systems. Hands-on experience operating production-grade environments, including EC2, networking, storage, monitoring, and cost management. Strong knowledge of Linux systems, containers, networking, storage, and distributed systems, with experience in orchestration or workload-scheduling technologies. Experience with infrastructure-as-code and delivery automation using tools such as Terraform or CloudFormation, configuration management, and CI/CD pipelines. Proficiency in Python and shell scripting, with sound software engineering practices for building maintainable automation, services, and platform integrations. Working knowledge of AI/ML development workflows and demonstrated ability to translate research requirements into platform designs, communicate technical trade-offs, manage delivery dependencies, and drive risk and control items to closure. Experience in one or more of the following domains: cloud and accelerated compute platforms (e.g., AWS, CPU/GPU infrastructure, autoscaling, workload scheduling, resource quotas, capacity planning); research environments and developer platforms (e.g., containers, Jupyter, package and dependency management, artifact repositories, experiment tracking, self-service tooling); AI workload enablement (e.g., LLM inference and evaluation, agentic workflows, retrieval pipelines, model-serving infrastructure, distributed processing); reliability and operations (e.g., observability, service-level objectives, incident response, resilience, backup and recovery, operational readiness) Preferred qualifications, capabilities, and skills Experience supporting AI/ML research or development teams and operating platforms for GPU-intensive, LLM, agentic, or other distributed AI workloads. Experience with Kubernetes or managed container platforms, GPU software stacks, cluster schedulers or distributed-compute frameworks such as Slurm or Ray, observability platforms, and cloud cost optimization. Experience delivering infrastructure in a regulated enterprise and working with cybersecurity, architecture, technology risk, and audit stakeholders. Familiarity with finance or financial use cases.