As a Senior DevOps Engineer, you will partner closely with Product, Engineering, and AI teams to define and evolve our cloud platform strategy. You will be responsible for designing secure, resilient, and scalable infrastructure that powers both modern applications and AI-driven solutions.
This role is instrumental in enabling production-ready AI systems through robust automation, strong observability, reliable deployment practices, and operational excellence across our technology ecosystem.
Key Responsibilities
-
Design, implement, and manage secure, scalable cloud infrastructure that supports business-critical applications and AI workloads.
-
Develop and enhance automation for infrastructure provisioning, CI/CD pipelines, and operational workflows to improve efficiency and reduce manual intervention.
-
Incorporate AI-powered capabilities into operational processes to improve system reliability, automate routine tasks, detect anomalies, and accelerate software delivery.
-
Ensure platform reliability through proactive monitoring, incident response, performance tuning, and capacity planning.
-
Champion DevOps best practices by promoting infrastructure as code, automated testing, repeatable deployments, comprehensive documentation, and operational excellence.
-
Work collaboratively with cross-functional teams to communicate architectural decisions, solve complex technical challenges, and support predictable software delivery.
-
Mentor engineers and provide technical guidance to strengthen platform engineering capabilities, improve DevOps maturity, and elevate engineering standards across the organization.
Requirements
Experience
-
6+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines.
Cloud & Infrastructure
-
Strong practical experience working with public cloud platforms, including AWS, GCP, and Azure.
-
Solid understanding of cloud networking, compute, storage, security, and cloud cost optimization strategies.
Containers & Infrastructure as Code
-
Extensive experience with containerization and orchestration technologies.
-
Strong expertise in Infrastructure as Code (IaC) using tools such as Terraform, Pulumi, or CloudFormation.
AI/ML Infrastructure
-
Experience supporting or deploying AI/ML platforms and workloads, including model serving, GPU-based infrastructure, or vector databases.
-
Candidates without direct AI infrastructure experience should demonstrate a strong understanding of the infrastructure requirements needed to operate AI systems in production.
Reliability Engineering
-
Proven experience building and operating highly available, scalable production environments.
-
Hands-on experience implementing zero-downtime deployment strategies such as Blue/Green deployments, Canary releases, Progressive Delivery, or Preview Environments.
Modern Delivery Practices
-
Experience implementing GitOps workflows using tools such as ArgoCD or Flux.
-
Familiarity with building self-service developer platforms that improve engineering productivity through environment automation and internal tooling.
Networking
-
Experience designing and managing API gateway solutions and edge routing across multi-cloud environments.
Platform Security
-
Strong understanding of infrastructure security, including identity and access management (IAM), secrets management, runtime protection, and platform hardening.
Observability
-
Practical experience implementing modern monitoring, logging, and observability solutions to maintain platform health and reliability.
Leadership & Communication
-
Excellent communication and stakeholder collaboration skills.
-
Ability to clearly explain technical concepts, influence engineering decisions, mentor team members, and promote engineering best practices.
Programming
-
Hands-on experience with modern programming languages such as Python, Node.js, or NestJS is highly desirable for building automation and extending platform capabilities.
Secure Engineering Practices
-
Strong understanding of secure software development practices, including secrets management, access control, and secure handling of sensitive information throughout the software development lifecycle.
-
Demonstrated commitment to protecting source code, technical documentation, and customer data while complying with company information security standards in a remote working environment.
-
Ability to work closely with QA, IT, and Information Security teams to remediate vulnerabilities, mitigate security risks, and support incident response activities when required.