Utwórz profil, aby pracodawcy mogli Cię znaleźć, otrzymywać lepiej dopasowane oferty pracy i szybciej aplikować.
  • Wyszukiwanie ofert pracy
  • Zapisane
  • Stwórz CV
    Nowe
  • Wynagrodzenia
  • Subskrypcje

Senior Site Reliability Engineer (AI/ML Platform)

20000 - 27000 zł
Pełny etat

Link Group

Zdalna
  • Praca zdalna

Who We're Looking For (Your Profile):

  • You are a true Site Reliability, Platform, or Infrastructure Engineer at heart, with a proven track record of managing complex, large-scale distributed systems.
  • Kubernetes is your natural habitat. You have deep, practical experience managing large-scale containerized environments and understand the complexities of orchestration under heavy load.
  • You speak the language of observability fluently, with hands-on experience using tools like Prometheus, Grafana, and distributed tracing systems to make systems transparent and understandable.
  • You are a strong programmer. You write clean, scalable automation scripts and infrastructure-as-code using Python or Go and tools like Terraform.
  • You have a genuine curiosity or, ideally, direct experience with the unique challenges of AI/ML infrastructure, such as model serving pipelines, inference engines, or managing GPU-accelerated workloads.
  • You are a problem-solver who takes full ownership of issues from start to finish. When you see a problem, you don't just fix it—you figure out how to prevent it from ever happening again.
  • You excel at collaboration and enjoy mentoring other engineers, helping them adopt SRE principles and build more reliable software from the ground up.

We are looking for a seasoned Site Reliability Engineer to join the team responsible for the backbone of our global AI/ML services. This isn't your typical SRE role. You won't just be maintaining systems; you'll be the guardian of a massive, distributed AI compute platform that processes workloads at an incredible scale. You will ensure that our AI models and GPU-powered infrastructure are not just fast, but fundamentally reliable, observable, and built to last.

If you are passionate about building and operating large-scale systems and are excited by the unique challenges of the AI/ML world, this is the role for you.

,[Build Bulletproof Observability: You will design and implement the "nervous system" for our AI platform. This means going beyond basic monitoring to build comprehensive observability with robust telemetry, insightful dashboards (Grafana), and intelligent alerting (Prometheus). You will define and track SLOs/SLIs to ensure our services meet their promises and drive improvements when they don't., Automate Everything: Your mantra is "if you have to do it twice, automate it." You will write clean, effective code in Python or Go to eliminate manual toil, create self-healing systems, and build sophisticated tooling that accelerates incident response and makes deployments safer., Own the Incident Response Lifecycle: When critical systems fail, you will be on the front lines. You will lead the charge in incident management, participate in a blameless on-call rotation, and conduct insightful post-mortems that lead to real, lasting improvements. You'll build the runbooks that others will rely on., Engineer a World-Class Deployment Pipeline: You will be a key contributor to our CI/CD ecosystem, building rock-solid integrations, automated safety checks, and seamless rollback capabilities to ensure that we can innovate at speed without sacrificing stability., Act as a Reliability Partner for Product Teams: You will work side-by-side with product engineers who are building the next generation of AI services. You will be their trusted advisor on reliability, helping shape their architecture and ensuring their products are operationally sound long before they hit production.] Requirements: Prometheus, Grafana, Python, Go, AI, Excel, SRE Additionally: Private healthcare, Sport subscription, Foreign languages classes, Life Insurance, Cafeteria system.
Oferta pracy dodana 4 dni temu