Platform Engineer
TLDR
Deploy and manage a platform across cloud-native and legacy systems for defence training, applying SRE practices to improve reliability.
Deploy and configure systems on MODCloud / D2S / OpenShift
Manage Kubernetes environments, networking and connectivity, as well as security-aligned configurations.
Build and maintain CI/CD pipelines, GitOps workflows, and environment configurations to ensure repeatable and reliable deployments.
Support running systems by monitoring health, diagnosing platform and infrastructure issues and resolving deployment failures.
Work closely with integration engineers, customers and modelling engineers.
Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools.
Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis.
This is not a pure cloud or greenfield platform role. You will be working across cloud-native services, legacy systems, and integrated simulation environments, ensuring they operate reliably as a single platform.
Strong systems and infrastructure mindset
Calm under pressure during outages or failures
Pragmatic and delivery-focused, with a bias toward keeping systems running.
Strong collaborator across engineering disciplines
Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems.
Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker.
Understanding of Distributed Systems in production.
Experience working within constrained or regulated environments (e.g. MODCloud, D2S, OpenShift) and adapting to their tooling and limitations.
Experience building and operating CI/CD pipelines with automated deployment workflows.
Familiarity with GitOps approaches and tools such as ArgoCD.
Strong understanding of Network fundamentals, Zero Trust solutions, service to service communications and distributed system connectivity.
Ability to diagnose issues across infrastructure, networking, and application layers.
Experience supporting or integrating legacy and non-cloud-native systems alongside modern infrastructure.
Experience in developing with Go or Python as well as shell scripting.
Experience applying Site Reliability Engineering (SRE) practices such as monitoring, alerting, incident response, and service reliability improvement.
Benefits
Flexible Work Hours
Flexible Working Hours - We’re not bound by the 9-to-5 model. Collaborate with your manager on determining a work schedule that suits you.
Free Meals & Snacks
Healthy Snacks & Drinks Provided - If you decide to come into the office, we have a range of snacks and drinks for you to enjoy.
Health Insurance
Private Medical & Dental Insurance - offered through Bupa.
compensation details
Honest about Compensation - We maintain a well defined salary range which a member of the Talent Team will discuss with you during the first call.
Paid Parental Leave
Enhanced Parental Leave - we’re proud to offer 26 weeks maternity leave and 4 weeks paternity leave at full pay.
Paid Time Off
Unlimited Paid Holiday - we value and support the need to maintain a strong work-life balance.
Remote-Friendly
Hybrid Working - we understand that a one-size-fits all approach doesn’t suit everyone.
Skyral designs and develops advanced modeling and simulation technology tailored for enterprise, defense, and national security sectors. Our solutions integrate data visualization, artificial intelligence, and digital twin capabilities to empower decision-makers with a comprehensive view of complex environments, facilitating informed decision-making for real-world challenges.