Senior devops engineer - aws
VeonIT · Pune, India
FULL TIMEpermanent
Job Description
We are seeking an experienced Senior Dev Ops Engineer (8+ years) to design, implement, and manage infrastructure and deployment pipelines across hybrid environments (on-prem data centers + cloud). The role requires expertise in AWS, Azure, Kubernetes, Ia C, CI/CD, and monitoring tools, along with strong troubleshooting and security practices.
Key Responsibilities
• Cloud & Infrastructure Management
Manage and support hybrid infrastructure including physical data centers, AWS (EC2, S3, RDS, Aurora, Route53, ACM, Lambda, Elasti Cache, Amazon MSK, EKS, ECS/Fargate, Elastic Beanstalk, Sage Maker/Bedrock, Secrets Manager, WAF, Cloud Watch), and Azure services. Design and implement Infrastructure-as-Code (Ia C) using Terraform, Pulumi, Cloud Formation with Sceptre, and Helm; manage Git Ops deployments with Argo CD/Flux. Configure and maintain Kubernetes clusters (EKS) — including Karpenter autoscaling, Bottlerocket nodes, Helm, and the Strimzi operator for Kafka — ensuring scalability and high availability.• CI/CD & Automation
Build, maintain, and optimize CI/CD pipelines with Jenkins, Git Lab CI, Bitbucket, and Git Hub (including the ongoing Jenkins to Git Lab CI migration), using Docker and Kaniko for container image builds. Automate deployment, scaling, and monitoring processes.• Monitoring & Troubleshooting
Set up and maintain observability using Grafana, Prometheus, Alertmanager, Loki, Cloud Watch, Kibana, and Open Search; manage on-call alerting and escalation via Pager Duty. Perform root cause analysis of performance issues, scaling bottlenecks, and high utilization scenarios. Participate in on-call support, ensuring timely resolution of incidents.• Security & Compliance
Manage SSL/TLS and code signing certificates (issuing, renewing, deploying) via AWS Certificate Manager (ACM), including DNS (CNAME) and email validation workflows. Implement best practices for infrastructure security, secrets and password management (AWS Secrets Manager, Hashi Corp Vault, 1 Password), and access control (IAM, AWS SSO/Okta, Organizations SCP). Support compliance and audit requirements through documentation and monitoring.• Collaboration & Knowledge Sharing
Work with cross-functional teams to improve infrastructure reliability and delivery processes. Maintain clear documentation in Confluence, track work in Jira, communicate via Slack, and contribute to knowledge sharing. Write root cause analysis (5 Whys) and incident post-mortems to prevent future issues.Required Skills & Experience
8+ years of experience in Dev Ops, Site Reliability Engineering, or related roles. Strong hands-on expertise with AWS services (EC2, S3, RDS, Aurora, Route53, ACM, Lambda, Elasti Cache, MSK, EKS, ECS/Fargate, Elastic Beanstalk, Secrets Manager, WAF, Cloud Watch) and Azure cloud. Solid experience with Infrastructure-as-Code tools (Terraform, Pulumi, Cloud Formation/Sceptre, Helm) and Git Ops (Argo CD/Flux). Proven experience with Kubernetes/EKS setup and operations (Helm, Strimzi, Karpenter, Bottlerocket, Docker). Proficiency in CI/CD pipelines (Jenkins, Git Lab CI, Bitbucket, Git Hub) and Git Ops (Argo CD/Flux). Strong knowledge of monitoring and observability tools (Grafana, Prometheus, Alertmanager, Loki, Cloud Watch, Kibana, Open Search) and incident/on-call tooling (Pager Duty). Scripting and automation skills in Python, Type Script, and Bash/shell. Good understanding of networking (IP, subnetting, VPC, security groups, ALB/ELB, Nginx, DNS/Route53, Open VPN), Linux/Windows Server administration, and troubleshooting. Hands-on experience with data streaming and CDC pipelines: Apache Kafka, Kafka Connect, Amazon MSK, and Debezium; familiarity with Databricks (Delta Lake, Asset Bundles, Unity Catalog) is a plus. Experience administering relational and in-memory data stores: My SQL, SQL Server, Maria DB, Amazon Aurora, and Redis/Elasti Cache (major/minor version upgrades and Blue/Green deployments). Configuration management and server automation with Salt Stack and/or Ansible. Experience managing SSL/code signing certificates (AWS ACM). Knowledge of security best practices in infrastructure and access control, and familiarity with compliance/audit frameworks (SOC 2, ISO 27001) and AWS security tooling (Trusted Advisor, IAM Access Analyzer).Location: Pune
If interested, kindly share your resume on info@
Details
| Company | VeonIT |
| Location | Pune, India |
| Type | FULL TIME |
| Niche | tech |
| Experience | permanent |
