Senior: PLN 260,400 - 352,200
Staff: PLN 350,700 - 474,400
Subject to alignment to the responsibilities and duties of the role - we currently have multiple positions available at both Senior and Staff level.
]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]" data-turn-id="request-WEB:e331ba19-dfb1-48a0-a51d-0973aacc5067-8" data-turn-id-container="request-WEB:e331ba19-dfb1-48a0-a51d-0973aacc5067-8" data-testid="conversation-turn-4" data-turn="assistant">
About the job
Make Graphcore’s observability platform dependable enough for the AI infrastructure it protects.
As a Senior QA Engineer in the Management & Observability team, you will validate Graphcore’s end-to-end telemetry platform. You will help ensure every metric, log, trace, dashboard, API and alert can be trusted at scale.
Your work will shape how quality is built into systems that monitor complex AI compute infrastructure. You will define test strategy, automate validation and expose issues before they reach production.
You will work across telemetry generation, collection, processing, storage, visualisation and alerting. You will build automated tests for distributed systems, production-like workloads and real operating conditions.
This role gives you ownership of quality in a platform where reliability, scale and engineering precision all matter.
The team and culture
You will work closely with telemetry and observability engineers from design through release. Quality is built early, through shared technical judgement and clear feedback.
The team moves by testing real assumptions, not waiting for perfect plans. You will have space to speak up, challenge designs and improve how releases are assessed.
Decisions are driven by evidence, root-cause analysis and practical engineering trade-offs. You will help set the standards that make Graphcore’s observability systems production-ready.
What we’re looking for
Experience designing automated test frameworks for distributed or infrastructure systems
Strong Python programming skills for functional, integration, system and regression testing
Experience testing Linux-based, cloud-native or containerised environments using Kubernetes or Docker
Understanding of CI/CD pipelines, API testing, networking and distributed system behaviour
Strong debugging and root-cause analysis skills for complex system-level issues
Familiarity with observability platforms, telemetry pipelines or time-series data
While we have outlined a set of requirements, we value transferable skills and diverse experiences. If you meet most of our essential criteria, we encourage you to submit an application and showcase how your background makes you a strong candidate.