Cogrion is building an Autonomous Data & AI Infrastructure Platform for modern enterprises.
We are looking for a hands-on DevOps Engineer who can build reliable, secure, automated cloud infrastructure for distributed data and AI workloads.
What You Will Work On
- Design, deploy, and operate Kubernetes-based platform infrastructure across cloud environments
- Automate infrastructure provisioning and application delivery using Terraform / OpenTofu, Helm, and GitOps
- Operate platform components such as Spark, Trino, Airflow, Jupyter, MLflow, and supporting services
- Build CI/CD pipelines, deployment automation, environment promotion, rollback, and release controls
- Implement observability using Prometheus, Grafana, Loki, OpenTelemetry, alerts, and operational dashboards
- Improve platform reliability, autoscaling, cost efficiency, security, backup, disaster recovery, and incident response
- Work with engineering teams to standardize secure runtime patterns for data, AI, and application workloads
What We Are Looking For
- Strong hands-on experience with Kubernetes and Docker in production environments
- Experience with at least one major cloud platform: AWS, Azure, GCP, or Alibaba Cloud
- Terraform / OpenTofu, Helm, GitOps, and infrastructure-as-code practices
- CI/CD experience with GitHub Actions, GitLab CI, Jenkins, Argo CD, or similar tools
- Linux, networking, DNS, ingress, load balancers, storage, IAM, and cloud security fundamentals
- Experience troubleshooting distributed systems, container workloads, and production incidents
- Good scripting skills in Python, Bash, or a similar language
Bonus Experience
- Spark on Kubernetes, Trino, Airflow, JupyterHub, Kafka, MLflow, or other data platform technologies
- Karpenter, Cluster Autoscaler, workload scheduling, GPU workloads, or advanced Kubernetes autoscaling
- Keycloak, OAuth / OIDC, workload identity, IAM roles, secrets management, and RBAC
- Multi-cloud or multi-region platform architecture
- FinOps, cloud cost optimization, SOC 2 / ISO 27001 controls, or platform security engineering
How We Work
- Own infrastructure and platform capabilities from design through production operations
- Automate repetitive work instead of relying on manual operational procedures
- Treat reliability, security, observability, and cost as product features
- Investigate root causes rather than applying temporary fixes
- Collaborate closely with product, backend, data, AI, and customer engineering teams
Why Cogrion?
What you should be comfortable owning: Architecture → Automation → Deployment → Observability → Reliability → Production
You will work on challenging infrastructure problems at the intersection of Kubernetes × Cloud × Data Infrastructure × AI. You will help build the platform layer that provisions compute, runs distributed workloads, secures enterprise environments, and keeps critical data and AI systems reliable at scale.
If you enjoy deep infrastructure engineering, production ownership, and building platforms that other engineers rely on, we would love to hear from you.
Interested in this role?
Send us your profile and a note about what you have built. We read every application.
Apply for this role