Purpose of the Job
To lead the implementation, optimization, and day-to-day operations of enterprise infrastructure platforms across on-premises and cloud environments.
This role combines hands-on technical execution with emerging leadership responsibility, ensuring high availability, security, and performance of infrastructure services while supporting strategic initiatives defined by management.
The Associate Manager acts as a technical owner and delivery driver, bridging architecture and operations, and ensuring infrastructure platforms are scalable, secure, and aligned with business needs.
Job Description
Customer
- Ensure high availability and reliability of infrastructure services, minimizing disruptions to business operations.
- Lead resolution of complex infrastructure incidents and escalations, acting as a technical subject matter expert.
- Drive root cause analysis (RCA) and implement preventive measures to enhance service stability.
Collaboration and Communication
- Collaborate with cross-functional teams (Security, Networking, Applications, DevOps) to ensure seamless infrastructure service delivery.
- Provide regular updates on infrastructure status, incidents, and ongoing initiatives to management.
- Mentor junior engineers and analysts, supporting skill development and technical growth.
- Act as a technical escalation point and guide best practices adoption across teams.
Data, Asset Management, and Analysis
- Monitor infrastructure performance, capacity, and availability, ensuring proactive optimization.
- Maintain accurate documentation for infrastructure architecture, configurations, and operational procedures.
- Support asset lifecycle management (hardware, software, cloud resources) with proper tracking and governance.
- Analyze system trends and recommend improvements to enhance performance and efficiency.
Financial Results
- Support cost optimization initiatives across infrastructure platforms (cloud and on-prem).
- Identify and implement efficiencies in resource utilization and provisioning.
- Contribute to infrastructure budget planning by providing technical input and cost forecasts.
Operations & Technical Ownership
- Infrastructure Management: Operate and maintain on-prem and cloud infrastructure ensuring performance, reliability, and security.
- Cloud & Platform Engineering: Implement scalable cloud solutions and support hybrid infrastructure environments. Hands on experience provisioning and implementing AWS infrastructure (EC2, EKS, S3, VPCs, etc) & landing zones.
- Network & Systems Support: Support network components, virtualization platforms, and compute/storage services.
- Security & Compliance: Enforce infrastructure security controls and support compliance requirements.
- Disaster Recovery & Resilience: Support design, provisioning and testing of disaster recovery and business continuity plans.
- Automation & IaC: Develop and maintain automation scripts and Infrastructure-as-Code (IaC) to improve provisioning and operations.
- Patch & Vulnerability Management: Ensure infrastructure components are regularly updated and secured.
- Monitoring & Observability: Implement and manage monitoring tools to ensure proactive incident detection and resolution.
- Vendor & ISP Coordination: Work with vendors and service providers to ensure service delivery and issue resolution.
Job Requirements - Experience and Education
- Bachelor’s degree in Computer Science, Information Technology, or related field.
- 6–8 years of experience in IT infrastructure, with strong hands-on technical expertise.
- Relevant certifications (AWS, GCP, VMware, etc.) are highly desirable.
- Proven experience managing enterprise environments across cloud (AWS, GCP & Azure) and on-prem infrastructure.
- Strong knowledge of:
- Virtualization technologies (VMware, Hyper-V)
- Networking (LAN/WAN, VPN, firewalls)
- Operating systems (Windows/Linux)
- Infrastructure security and compliance standards
- Experience with automation, scripting, and Infrastructure-as-Code (IaC) tools. Terraform preferred.
- Good understanding and hands on expertise with high availability, disaster recovery, and performance optimization strategies.