Manage and support Linux servers, Kubernetes clusters, and containerized workloads.
Administer and optimize GCP resources and cloud services.
Support platform services including Elasticsearch, Redis, RabbitMQ, and associated components.
Monitor platform health, integrations, logs, alerts, and overall system performance.
Handle critical incidents, escalations, troubleshooting, and root cause analysis (RCA).
Support deployments, upgrades, configuration changes, and platform enhancements.
Monitor integration health and proactively resolve issues and anomalies.
Coordinate with internal cloud, infrastructure, and engineering teams.
Maintain operational documentation, SOPs, and technical records.
Support automation initiatives and continuous service improvements.
Ensure platform availability, reliability, scalability, and operational efficiency.
Desired Candidate Profile
3-5+ years of experience in platform engineering, cloud operations, system administration, or DevOps environments.
Strong experience with Linux, Kubernetes (GKE), and Google Cloud Platform (GCP).
Experience supporting high-availability production environments, platform services, and integrations.
Exposure to monitoring, observability, automation, and incident management processes.
Skills
Linux (Ubuntu), Bash/Python
Kubernetes (GKE), Docker
Google Cloud Platform (Compute, IAM, Networking)
ELK Stack (Elasticsearch, Logging)
Redis, RabbitMQ
REST & GraphQL APIs
Prometheus, Grafana, Loki
YAML, Git
Ticketing tools (Jira/ServiceNow)