← Browse all jobs

C

Senior Data Engineer

Commotion · India

FULL TIME

Job Description

Commotion builds an AI operating system for large enterprises, and at the centre of it is a context graph: a live model of an organisation’s entities, the relationships between them, and the decisions taken across them, assembled from the systems that already run the business.

Most enterprise AI stops at retrieval over documents. We think the layer that matters next is structural — resolved entities, explicit relationships, and a record of every decision and the reasoning behind it. A dashboard tells you what happened. A context graph lets an agent work out what to do about it, and show its working.

We build this inside our clients’ own environments, where the data is real, messy and consequential. We are backed by Tata Communications, and we are early enough that the people joining now shape how this gets built.

About this role

The data a client hands us is never in the shape we need, and it is never as clean as they believe. This role owns the path from their source systems to something downstream can be trusted with, and keeps it correct as those systems change underneath us.

Infrastructure, deployment and monitoring stacks are owned by our platform team. This role is about the data.

What you’ll do

  • Build the core ingestion framework, so that adding a new source is one integration and everything else — scheduling, retries, reconciliation, lineage, snapshotting, monitoring — comes for free. This is the single highest-leverage thing in the role.
  • Integrate with client source systems: ERPs, CRMs, warehouses and lakehouses, on-prem databases, APIs, file feeds and object storage. Enterprise authentication, rate limits, pagination, and the undocumented behaviour you only find in production.
  • Run this at real scale. Multi-terabyte data lakes on the source side, graphs with billions of nodes and edges on the target side, refreshed on a daily or hourly cycle.
  • Build ingestion pipelines that stay correct under that load: incremental loads, idempotency, ordering, checkpointing, resumable failure recovery, and backfills that do not double count.
  • Own snapshotting and versioning of loaded data, so a bad load can be rolled back, a previous state can be reproduced, and two versions can be compared.
  • Publish new data without disturbing live applications reading it. Atomic swaps, versioned reads, and refreshes that never leave an agent querying a half-written graph.
  • Treat data correctness as a deliverable in its own right: reconciliation against source, detection of schema and semantic drift, and alerting that fires before a client notices anything.
  • Own performance and cost at volume. Know when a pipeline is slow because of the data, the query, the partitioning, or the cluster.
  • Set the engineering bar for the team, and review work that is not your own.

What we’re looking for

  • 6 or more years building and operating production data pipelines.
  • You have worked at terabyte scale and above, and can talk concretely about what changed in your design because of it — partitioning, compaction, skew, shuffle cost, incremental strategy.
  • Strong Python and SQL. Spark at volume. A workflow orchestrator you have run in production, not just configured.
  • Deep experience integrating messy enterprise systems. Undocumented APIs, broken pagination, soft deletes, inconsistent timezones, duplicate keys, silently truncated exports. You have been burned and you know where to look.
  • A working grasp of the distributed data problems that actually bite: idempotency, exactly-once versus at-least-once, late and out-of-order data, checkpointing and resumability, backfill without double counting.
  • You have built or extended a framework that other engineers then used, and you can explain the abstractions you chose and the ones you regretted.
  • Versioned or immutable data patterns: snapshots, time travel, atomic publishes, rollback. You understand why a live reader must never see a partial write.
  • Data quality engineering as a habit, not a phase. Reconciliation checks, contracts, tests on data as well as on code.
  • Cloud data platforms and object storage at scale, on at least one of AWS, Azure or GCP.
  • You have owned something in production that other people depended on, and you were the one paged when it broke.

How we work

  • We are not asking you to invent a framework from scratch. We are asking you to know what to reach for, where it breaks, and how you found that out. Bring opinions about tools you have run in production, including ones you would never use again.
  • Every engineer here works with Claude and agentic coding tools daily, for exploration, transformation code, test data and analysis scaffolding. Our delivery pace assumes it. Be ready to describe how these tools changed your workflow and where you have learned not to trust them.
  • This is client-facing work on client sites in India. Expect travel, and expect to sit with the client’s own data owners.



Nice to have

  • Graph databases in production — modelling and operating them, not only querying.
  • Change data capture and streaming tooling.
  • On-prem or air-gapped delivery, and the security conversations that come with it.
  • Warehouse and lakehouse internals deep enough to tune cost, not just correctness.
  • Experience building an integration framework that others then used.

Details

CompanyCommotion
LocationIndia
TypeFULL TIME
Nichegeneral

Similar Jobs

S

Customer Support Representative

Skillinabox

I

Workday Consultant

Innodata Inc.

S

Lead Backend Engineer

SuperOps

S

Associate Manager - Business Development

Salt Consult

C

Area Sales Manager

Concept Medical