{bc}
greenhouse

Principal Software Engineer (workflow / platform)

Cognite
Bengaluru, IND
Full-time
Senior · 10+ years experience
Onsite
Discovered 1 weeks ago
KotlinJavaPythonFastAPIKubernetesAzure
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

KotlinJavaPython
Smart Apply

Full Job Posting

What Cognite is: Relentless to achieve

Cognite operates at the forefront of industrial digitalization, building AI , and data solutions that solve the world’s hardest, highest-impact problems. With unmatched industrial heritage and a comprehensive suite of AI capabilities, including low-code AI agents, Cognite accelerates the digital transformation to drive operational improvements.

We thrive in challenges. We challenge assumptions. We execute with speed and ownership. If you view obstacles as signals to step forward - not backwards - you’ll feel right at home here.

Our Moonshot is bold: Unlock $100B in customer value by 2035, and redefine how global industry works. Join us in this venture where AI and data meet ingenuity, and together, we will forge the path to a smarter, more connected industrial future.

How you’ll demonstrate Ownership

Platform Ownership: Design, build, and operate the core serverless execution engine and Workflows orchestration layer that serve as foundational primitives for CDF’s AI and automation capabilities.

Reliability Engineering: Own uptime, latency SLOs, and incident response for platform services ensuring Functions execute deterministically and Workflows progress without data loss or silent failures.

Scalability: Architect for multi-tenant & multi-cloud, high-throughput workloads. Design scheduling, queueing, and retry mechanisms that degrade gracefully under pressure.

API Design: Define and evolve clean API-first architecture, versioned REST and event-driven APIs that downstream engineering teams and external customers depend on.

Observability: Instrument services with distributed tracing, structured logging, and alerting (Open-telemetry / Prometheus / Grafana / Honeycomb stack) so failures surface before customers notice.

CI/CD & Testing: Champion test automation - unit, integration, and smoke tests and maintain deployment pipelines that ship to production with confidence.

Performance: Profile and resolve bottlenecks in execution throughput, cold-start latencies, and cross-service call chains driving a “snappy” platform experience for industrial workloads.

Cost Efficiency (Bonus): Model compute and storage costs for functions execution; identify and implement optimizations that reduce cloud spend without sacrificing reliability.

The Impact you bring to Cognite

10+ years of Engineering: Proven track record building and operating production backend services at scale.

Expertise: Deep mastery of JVM languages (Kotlin preferred, Java acceptable), Python(FastAPI), distributed systems patterns, and cloud-native service design (Kubernetes, Azure, GCP, AWS, Private cloud).

Workflow & Orchestration: Hands-on experience with workflow engines (Conductor, Apache Airflow, or equivalent) and event-driven architectures (Kafka, Pub/Sub).

Data & Storage: Comfortable working with relational databases (PostgreSQL) & non-relational databases, object storage(Data-lakes), and caching layers (Redis) in multi-tenant environments.

Observability Stack: Practical experience with Open-telemetry, Prometheus, and Grafana for instrumentation and operational insight.

ML Platform Exposure: experience supporting ML workloads & notebooks in production, whether through job scheduling, resource management, experiment tracking integration, or model serving infrastructure.

Contextualisation Domain (Bonus): Familiarity with industrial knowledge graph construction, entity resolution, or NLP/CV pipelines as they relate to industrial asset data is a strong differentiator.

Full-Stack Awareness (Bonus): Familiarity with React or TypeScript is a plus for consuming and dogfooding your own platform’s developer tooling.

The Platform Thinking Spirit: A passion for building composable, well-documented, and automated platform systems that empower other engineers including ML engineers to build faster.

ML Platform & Contextualisation

ML Workload Support: Build and extend platform primitives compute scheduling, environment management, and secrets handling, that enable ML engineers to run model training, fine-tuning, and batch inference jobs reliably.

Contextualisation Pipelines: Support the engineering infrastructure behind Cognite’s Contextualisation capabilities (entity matching, asset hierarchy inference, P&ID parsing) by ensuring the platform can orchestrate long-running, GPU-aware, and data-intensive ML workflows without manual intervention.

Vector & Embedding Infrastructure (Bonus): Familiarity with serving or storing vector embeddings to support semantic search and RAG-based contextualisation use cases.

Model Lifecycle Awareness: Understand model versioning, A/B experiment tracking, and the boundary between platform concerns and ML framework concerns, so the platform stays lean while ML teams stay unblocked.

Impact 2025

Cognite's Industrial AI: Moonshot

We’re globally recognized domain experts with an international presence that spans Phoenix, Houston, Oslo Tokyo, Bengaluru, and Abu Dhabi.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Cognite