Utwórz profil, aby pracodawcy mogli Cię znaleźć, otrzymywać lepiej dopasowane oferty pracy i szybciej aplikować.
  • Wyszukiwanie ofert pracy
  • Zapisane
  • Stwórz CV
    Nowe
  • Wynagrodzenia
  • Subskrypcje

Senior HPC DevOps Engineer

EPAM Systems

Katowice
  • Praca zdalna

To stanowisko wymaga obecności na miejscu. Zobacz podobne oferty poniżej.

We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform. Responsibilities Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure frictionAdvise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloadsMaintain automated pipelines for infrastructure provisioning and platform service deploymentsResolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotasCollaborate with developer experience teams to improve documentationCollaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or PrometheusManage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelinesEnsure compute availability through capacity planning and reservation managementDeploy containerized environments tuned for HPC and GPU pass-throughDeploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions Requirements 5+ years of experience in HPC or DevOps engineering rolesKnowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirectProven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/InferentiaHands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota managementExperience in deployment of containerized environments using Apptainer/Singularity, Docker, or EnrootUnderstanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for LustreHands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecksExperience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)Proficiency in English at a B2+ level We offer We gather like-minded people:Top tech minds driving innovation in AI, cloud and digital platform modernizationSupportive team and agile, startup-like cultureHybrid by design mode and opportunity to work remotely within PolandChance to work abroad for up to 60 days annuallyBusiness-driven relocation opportunitiesWe provide growth opportunities:Career development programsThought leadership, mentoring, soft skills and well-being programsCertification (Anthropic, Gemini, GCP, Azure, AWS)English classesWe cover it all:Stable payParticipation in the Employee Stock Purchase Plan with a 15% discountBenefits package (health insurance, multisport, shopping vouchers)Referral bonuses up to $2,000Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and moreCorporate, social and well-being eventsPlease, note:Benefits listed above are available to employees onlyWe are open for working with Contractors. Terms of B2B cooperation agreements are agreed individuallyWe will reach out to selected candidates exclusively EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

Oferta pracy dodana 9 godzin temu