Site Reliability Engineer
Il y a 20 heures
Luxembourg
NTT DATA Europe & Latam
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Site Reliability Engineer (SRE)
- European Institution Are you looking for the next step in your career? Join NTT DATA Would you like to pursue your career in Public Institutions? We are currently looking for a Site Reliability Engineer to join our growing client in Luxembourg
Key Responsibilities
• Build, deploy, and operate secure, scalable and resilient on-prem automation and orchestration platforms using technologies such as n8n, Cycloid, Ansible and Kubernetes/OpenShift ;
• Design and maintain workflow/orchestration capabilities, including triggers/webhooks, retries, timeouts, error handling, secrets/credentials, versioning, environment promotion, reusable components and governance ;
• Automate infrastructure provisioning and configuration and develop reusable platform building blocks for internal teams ;
• Develop automated platform tests and define/test Disaster Recovery procedures ;
• Implement and improve observability, including monitoring, logging, alerting, tracing, dashboards and platform health ;
• Ensure platform reliability, scalability, performance and security, troubleshoot production issues, and support on-call operations ;
• Standardize and support the onboarding of internal teams, working with security and application teams to ensure compliance with best practices and regulatory requirements. Profile & Experience
• 5+ years of SRE experience, ideally supporting internal platforms ;
• Bachelor's or Master's Degree in Computer Sciences ;
• Hands-on experience with workflow/orchestration platforms such as n8n or comparable tools, including resilient workflow execution, reusable components, secrets management, promotion/versioning and multi-team governance ;
• Strong experience with Docker, Kubernetes and Kubernetes Operators; OpenShift is a plus ;
• Experience operating IT services at scale, including production support and on-call participation ;
• Strong scripting/automation skills using Python, Bash and/or Go ;
• Knowledge of GitOps principles and tooling; GitLab is a plus ;
• Good understanding of Platform Engineering, infrastructure automation, observability, resilience and Disaster Recovery practices ;
• Strong troubleshooting and problem-solving skills, including the ability to perform under pressure ;
• Ability to work independently and across multiple teams, with experience in Agile environments and basic ITIL knowledge.
Requirements:
• This role requires working in shifts (7:00-16:00 or 11:00-20:00)
• On-call: mandatory rotation covering 7 consecutive days, including weekends What we have to offer :
• Flexibility : chose to work as a full time employee or as a freelancer with attractive conditions in both cases
• Comfortable working
conditions:
hybrid or remote set-up (from Europe only)
• Extra perks if you work as a full time employee.
- European Institution Are you looking for the next step in your career? Join NTT DATA Would you like to pursue your career in Public Institutions? We are currently looking for a Site Reliability Engineer to join our growing client in Luxembourg
Key Responsibilities
• Build, deploy, and operate secure, scalable and resilient on-prem automation and orchestration platforms using technologies such as n8n, Cycloid, Ansible and Kubernetes/OpenShift ;
• Design and maintain workflow/orchestration capabilities, including triggers/webhooks, retries, timeouts, error handling, secrets/credentials, versioning, environment promotion, reusable components and governance ;
• Automate infrastructure provisioning and configuration and develop reusable platform building blocks for internal teams ;
• Develop automated platform tests and define/test Disaster Recovery procedures ;
• Implement and improve observability, including monitoring, logging, alerting, tracing, dashboards and platform health ;
• Ensure platform reliability, scalability, performance and security, troubleshoot production issues, and support on-call operations ;
• Standardize and support the onboarding of internal teams, working with security and application teams to ensure compliance with best practices and regulatory requirements. Profile & Experience
• 5+ years of SRE experience, ideally supporting internal platforms ;
• Bachelor's or Master's Degree in Computer Sciences ;
• Hands-on experience with workflow/orchestration platforms such as n8n or comparable tools, including resilient workflow execution, reusable components, secrets management, promotion/versioning and multi-team governance ;
• Strong experience with Docker, Kubernetes and Kubernetes Operators; OpenShift is a plus ;
• Experience operating IT services at scale, including production support and on-call participation ;
• Strong scripting/automation skills using Python, Bash and/or Go ;
• Knowledge of GitOps principles and tooling; GitLab is a plus ;
• Good understanding of Platform Engineering, infrastructure automation, observability, resilience and Disaster Recovery practices ;
• Strong troubleshooting and problem-solving skills, including the ability to perform under pressure ;
• Ability to work independently and across multiple teams, with experience in Agile environments and basic ITIL knowledge.
Requirements:
• This role requires working in shifts (7:00-16:00 or 11:00-20:00)
• On-call: mandatory rotation covering 7 consecutive days, including weekends What we have to offer :
• Flexibility : chose to work as a full time employee or as a freelancer with attractive conditions in both cases
• Comfortable working
conditions:
hybrid or remote set-up (from Europe only)
• Extra perks if you work as a full time employee.