Cogrion is building an Autonomous Data & AI Infrastructure Platform for modern enterprises.
We are looking for a strong Data Platform Engineer who can design and build scalable data infrastructure, distributed processing capabilities, and reliable platform services for enterprise workloads.
What You Will Work On
- Build and improve core data platform capabilities around Spark, Trino, lakehouse storage, orchestration, and metadata
- Design scalable batch, streaming, CDC, and SQL processing patterns for enterprise workloads
- Build platform abstractions that simplify compute, ingestion, optimization, governance, and data operations for users
- Improve performance, concurrency, reliability, resource management, and cost efficiency of distributed data workloads
- Develop APIs, services, libraries, and automation around data engines and platform workflows
- Solve platform-level problems such as small files, data skew, compaction, schema evolution, transaction conflicts, and workload isolation
- Work across Kubernetes and cloud infrastructure to operate data engines reliably at scale
What We Are Looking For
- Strong experience with Apache Spark and distributed data processing
- Strong SQL skills and experience with Trino, Presto, or similar distributed query engines
- Experience with Delta Lake, Apache Iceberg, Apache Hudi, or modern lakehouse architectures
- Hands-on experience with batch and streaming data pipelines, preferably using Kafka or similar systems
- Strong Python, Java, or Scala programming skills
- Good understanding of data partitioning, file formats, query optimization, concurrency, and distributed systems
- Experience working with object storage such as S3, OSS, ADLS, or GCS
Bonus Experience
- Kubernetes-based Spark or Trino deployments and workload scheduling
- Airflow, Jupyter, MLflow, Hive Metastore, DataHub, or related platform technologies
- CDC technologies, Kafka Connect, Debezium, dlt, or streaming ingestion frameworks
- Metadata, lineage, data quality, governance, or semantic layer technologies
- Cloud infrastructure, Terraform / OpenTofu, Helm, observability, and platform operations
- Experience building an internal data platform, developer platform, or enterprise SaaS product
How We Work
- Think in terms of reusable platform capabilities rather than one-off pipelines
- Own engineering problems from architecture and implementation through production behavior
- Profile and diagnose distributed workloads using evidence, metrics, and execution internals
- Design for scale, failure recovery, upgradeability, maintainability, and predictable economics
- Collaborate across platform, DevOps, backend, AI, and product engineering teams
Why Cogrion?
What you should be comfortable owning: Problem → Architecture → Engine Integration → Performance → Reliability → Production
You will work on difficult engineering problems at the intersection of Data Infrastructure × Distributed Systems × Cloud × AI. Rather than building isolated ETL jobs, you will help build reusable infrastructure that allows enterprises to ingest, process, query, govern, and operationalize data at scale.
If you enjoy understanding how data engines work internally and turning complex infrastructure into simple product experiences, we would love to hear from you.
Interested in this role?
Send us your profile and a note about what you have built. We read every application.
Apply for this role