Utwórz profil, aby pracodawcy mogli Cię znaleźć, otrzymywać lepiej dopasowane oferty pracy i szybciej aplikować.
  • Wyszukiwanie ofert pracy
  • Zapisane
  • Stwórz CV
    Nowe
  • Wynagrodzenia
  • Subskrypcje

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

T-Mobile

Expected, Kubernetes, OpenShift, vLLM, Paged Attention, Continuous Batching, NVIDIA GPU, CUDA, NVIDIA Container Toolkit, Prometheus, Grafana, OpenTelemetry, ELK Stack, Python, Bash, Linux, GitLab CI, Jenkins, ArgoCD, Infrastructure as Code (IaC), MLOps, AI Infrastructure, LLM, DevOps, SRE, Platform Engineering

Your responsibilities, Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure., Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants., Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3., Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency., Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing., Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates., Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements., Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics., Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns., Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services., Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices., Maintain audit logging capabilities to support compliance, security investigations, and operational governance., Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services., Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.

5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations., At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms., Strong experience with Kubernetes and OpenShift administration in production environments., Proven experience deploying and operating vLLM-based inference platforms in production., Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques., Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack., Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms., Strong Python programming skills with experience developing automation and operational tooling., Experience with Bash scripting and Linux systems administration., Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices., Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.

What we offer, Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation., You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.

Benefits, sharing the costs of sports activities, private medical care, sharing the costs of professional training & courses, life insurance, remote work opportunities, flexible working time, corporate products and services at discounted prices, mobile phone available for private use, no dress code, parking space for employees, extra social benefits, sharing the costs of tickets to the movies, theater, holiday funds, birthday celebration, sharing the costs of a streaming platform subscription, employee referral program, charity initiatives, extra leave, platforma benefitowa

Recruitment stages, Prześlij swoje CV, Spotkaj się z przyszłym liderem/ liderką zespołu, Witaj w T-Mobile :)

T-Mobile, We are a technology company, and our goal is to create innovative solutions for individual and business clients., , At T-Mobile, we all live in a magenta world! This color is close to our hearts and means faith in the success of undertaken actions, self-confidence, and endurance., , That’s who we are as a team., , At #MagentaTeam , we focus on exchanging experiences, agile work, and quick adaptation to changes! #MagentaTeam is, above all, a mix of different competencies, experiences, personalities, temperaments, and views. And this diversity is our greatest strength.

This is how we work,
Oferta pracy dodana 2 dni temu