← Browse all jobs

T

Senior data engineer

Teqfocus · Pune, India

FULL TIMEpermanent

Job Description

About the Opportunity:

Teqfocus is hiring a Senior Data Engineer for a confidential client engagement focused on building the data foundation behind a modern, domain-specific intelligence platform. Client details will be shared at the appropriate stage of the hiring process. This is a full-time hire and the need is immediate.

This is a hands-on engineering role for someone who wants to build durable, domain-specific data assets: ingestion pipelines, layered transformations, AI-ready datasets, vector and embedding flows, observability, data quality, and governed data products that analytics, ML, RAG, semantic search, and agentic workloads can rely on.

About Teqfocus:

Teqfocus is a Data and AI company that helps enterprises build and scale the Agentic Enterprise — connecting the data foundation, the intelligence layer, and the applications teams run their businesses on into a single, coherent architecture.

As a Salesforce Summit Partner, we work in Healthcare, Life Sciences, Financial Services, Insurance, and Hi-Tech — industries where getting AI right matters and the tolerance for failure is low. Our delivery model is senior-heavy by design. The same team that scopes an engagement ships it to production. There are no junior shadow teams, no sub-contractors brought in when the problem gets hard, and no hand-off to a separate managed services vendor when the project closes.

We run an internal innovation practice that builds the reusable patterns, pre-tested architectures, and industry-specific frameworks that give every Teqfocus client a head start — and give every Teqfocus practitioner work that compounds over time.

We are a diverse, immigrant-entrepreneurial company. We were built on the conviction that the best outcomes come from people who take genuine ownership of what they build — and are trusted to do so.

Location - Hybrid - Pune/Ranchi

Key Responsibilities:

Pipeline Development and Operation

Build and operate ingestion, ELT/ETL, and orchestration pipelines that move data from Mongo DB Atlas and other sources into analytical and AI-serving layers. Implement layered, medallion-style transformations with idempotent, backfillable, incrementally loaded jobs. Apply deduplication, normalization, and validation so downstream data is high-quality and trustworthy. Modernize legacy or homegrown data flows through incremental strangler-fig migrations that keep production stable.

AI-Ready Data Delivery

Build embeddings and vector pipelines, along with feature-ready and retrieval-ready datasets for RAG, semantic search, and agentic workloads. Make production data AI-ready in practice: well-structured, lineage-tracked, governed, and retrieval-friendly. Implement real-time and change-data-capture flows from Mongo DB using Change Streams / CDC where freshness is required.

Architecture and Data Contracts

Implement canonical data models, schemas, and data contracts defined by the Data Architect, enforced in-repo so teams build against stable definitions. Exercise sound persistence judgment in execution, landing data in the right store: document / No SQL, vector, analytical, or lakehouse, according to architectural direction. Contribute to build-vs-buy decisions through prototypes using proven, industry-standard tooling over unnecessary custom development.

Quality, Reliability and Observability

Establish testing, data-quality, and lineage checks for owned pipelines, with clear alerting and runbooks. Instrument pipeline observability across freshness, volume, schema drift, failures, and cost so issues are caught before consumers feel them. Use AI-assisted development tools such as Claude Code, Git Hub Copilot, and Cursor as force multipliers for transformation logic, query tuning, migration scripting, test generation, and documentation.

Collaboration

Partner with database engineering on extracting from and protecting the production Mongo DB store. Partner with the Data Architect on implementing target-state patterns and surfacing what is hard to build. Partner with ML, AI, analytics, and application engineers on the data they consume, shaping and governing it so it is safe and ready to build on.

Required Skills and Experience:

8+ years of hands-on data engineering experience building and operating production data pipelines at scale. Strong programming and data skills in Python and SQL, with solid software engineering fundamentals: version control, testing, CI, production code ownership, maintainability, and operational support. Hands-on Mongo DB experience at production scale. Mongo DB Atlas is ideal. Must understand document modeling, aggregation framework, Change Streams / CDC, and extraction from document stores into analytical and AI-serving layers. This role is No SQL / Mongo DB-focused, not relational-first. Demonstrated experience with ELT/ETL pipeline design, transformation frameworks such as dbt or equivalents, and orchestration tools such as Airflow, Dagster, or Azure Data Factory. Experience building on cloud-native data platforms and lake, lakehouse, or warehouse architectures using layered medallion-style modeling. Hands-on experience preparing data for AI/ML or analytical consumers, including embeddings / vector pipelines, RAG-ready datasets, feature-ready datasets, deduplication, normalization, and validation. Familiarity with vector search and embeddings in production, ideally Mongo DB Atlas Vector Search or equivalent. Demonstrated use of AI-assisted development tools such as Claude Code, Copilot, or Cursor for data and pipeline work. Strong grasp of data quality, testing, lineage, data contracts, pipeline observability, alerting, and operational runbooks. Comfortable working in a complex, specialized domain. Experience with data-intensive product environments, operational workflows, BIM, CAD, multimodal, or unstructured data is a plus; appetite to learn the domain is required.

Preferred / Nice-to-Have Experience:

Azure data ecosystem experience: Azure Data Factory, Synapse Analytics, Azure Functions, Event Grid, Event Hubs, or related services. Lakehouse platforms such as Databricks or Snowflake, or open table formats such as Iceberg, Delta, or Hudi. Feature stores such as Feast or equivalent. Streaming or event-driven processing using Kafka, Event Hubs, or Spark Structured Streaming. CDC and cross-engine synchronization using Mongo DB Change Streams, Debezium, or equivalent. Experience with geometric, BIM, CAD, multimodal, operational, or unstructured source data. Knowledge graph, ontology, semantic-layer, or governed semantic search exposure. Data governance for AI or agent access to production data, including query-cost controls, read-path safety, lineage, and audit. SOC 2 and data classification awareness.

AI-Native Engineering Expectations:

Teqfocus values senior engineers who use AI-assisted development responsibly as part of a real engineering loop. The right candidate can use AI to move faster while retaining ownership of correctness, security, maintainability, and production quality.

Uses tools such as Claude Code, Cursor, Git Hub Copilot, or similar assistants for planning, coding, refactoring, test creation, debugging, documentation, and migration support. Breaks work into small iterations, validates generated code, checks for hallucinated APIs or weak tests, and can explain why a chosen approach is correct. Uses AI as a thinking partner for implementation sequencing, query tuning, transformation logic, and data-quality rule development without outsourcing judgment.

What Success Looks Like in the First Year :

Pipelines feeding the intelligence platform are built, reliable, observable, and production-ready; data lands fresh, clean, and on schedule. Curated AI-ready datasets and embeddings/vector flows are in production, enabling AI/ML and agentic work to build on a trusted substrate. The Data Architect's canonical model and contracts are implemented and enforced in pipelines, not just documented on paper. Legacy and homegrown data flows are replaced incrementally with proven, maintainable tooling, without big-bang disruption. Data quality and lineage checks catch issues before downstream consumers do.

Please Apply If You Are:

A hands-on senior data engineer who wants to build production data systems that power AI and analytics at scale. Deeply comfortable with Mongo DB / No SQL data engineering and excited to transform operational data into governed, AI-ready data products. Strong in ownership, written communication, debugging, operational discipline, and cross-functional collaboration. Interested in an immediate full-time opportunity through Teqfocus for a high-impact confidential client engagement.

What We Offer:

Competitive CTC with performance-based incentives Transparent compensation and standard benefits (PF, Gratuity, HRA, etc.) Medical, accident, and life insurance coverage Wellness and EAP support Paid leaves and public holidays Hybrid/remote flexibility based on role Direct client exposure across US & Canada Opportunities in architecture, innovation, and global projects Sponsored certifications and continuous learning Fast-paced, ownership-driven culture Flat hierarchy with direct leadership access Collaborative global team environment

The agentic enterprise is not a future state. It is being built right now, on live production systems, for real clients. If you want to be one of the practitioners building it — this is the right place. Apply Now!!!

Details

CompanyTeqfocus
LocationPune, India
TypeFULL TIME
Nichetech
Experiencepermanent

Similar Jobs

P

Registered Veterinary Technician

Portland Vet

P

Veterinary Technician - General Practice

Portland Vet

H

Physical Therapist Assistant Per Diem

Hebrew SeniorLife

U

Radiologic Technician

UCHealth

P

Veterinary Technician

Portland Veterinary Emergency and Specialty Care