← Browse all jobs

T

Gcp data engineer postgres python pyspark developer

Tata Consultancy Services · Hyderabad g.p.o., India

FULL TIMEpermanent

Job Description

Information Details

1 Role** GCP Data Engineer Postgres Python Pyspark Developer (Big Query, Cloud Storage, Dataproc, Airflow)

2 Required Technical Skill Set** GCP Data Engineer to design, build, and optimize scalable data pipelines and analytics solutions using Big Query, Cloud Storage, Dataproc, and Airflow.

3 No of Requirements** 5

4 Desired Experience Range** 7+ Years

5 Location of Requirement HYDERABAD

6 Immediate Joiners Needed YES

Desired Competencies (Technical/Behavioral Competency)

Must-Have** (Ideally should not be more than 3-5)

GCP Services: Big Query, Cloud Storage, Dataproc, Cloud Composer (managed Airflow) or self-managed Airflow. Airflow: Strong experience in DAG creation, operators/hooks, scheduling, backfilling, retry strategies, and CI/CD for DAG deployments. Programming: Proficiency in Postgres Programming including Tables, Triggers, Views, Stored procedures, Python and Pyspark (Py Spark, Airflow DAGs), SQL (advanced Big Query SQL). Data Modeling: Dimensional modeling (Star/Snowflake), data vault basics, and schema design for analytics. ·

Performance Tuning: Big Query partitioning/clustering, predicate pushdown, job stats review, Dataproc executor tuning. ·

Version Control & CI/CD: Git, branching strategies, pipelines for deploying Airflow DAGs and config.

· Operational Excellence: Monitoring with Stackdriver/Cloud Logging, debugging pipeline failures, and root-cause analysis. · involves end-to-end ownership of data ingestion, transformation, orchestration, and performance tuning for batch and near real-time workflows.

Good-to-Have

Streaming: Pub/Sub, Dataflow (Apache Beam) for near real-time pipelines. Orchestration Patterns: Event-driven pipelines, dependency management, and cross-environment promotion. Data Governance: Catalog/lineage tools (e.g., Data Catalog), PII handling, row-level security, column-level encryption. Containers & Infra: Docker, Terraform for Ia C on GCP; Kubernetes concepts. BI Integration: Experience integrating with Looker, Tableau, or Power BI. Certifications: Google Professional Data Engineer / Cloud Architect.

Responsibility of / Expectations from the Role

1 Data Pipeline Development: Build robust ETL/ELT pipelines using Apache Airflow (DAG creation, scheduling, monitoring) to orchestrate data workflows across GCP services.

2 Data Warehousing: Design and optimize Big Query schemas, partitioning/clustering strategies, materialized views, and query performance tuning.

3 Data Processing: Implement scalable data processing using Dataproc (Spark/Hive), including job configuration, optimization, and cost control.

4 Data Ingestion & Storage: Manage ingestion from diverse sources (APIs, files, streaming) and design storage strategies using Cloud Storage (lifecycle policies, tiers, security).

5 Quality & Observability: Implement data validation (e.g., Great Expectations or custom checks), logging/alerting, and SLA monitoring for pipelines.

6 Security & Governance: Apply IAM, service accounts, VPC SC, CMEK, and access policies across GCP resources; ensure compliance with data governance standards.

7 Cost & Performance: Optimize queries, cluster usage, and storage to balance cost/performance; leverage reservation/flex slots and job-level optimizations.

8 Collaboration: Work closely with analytics, product, and business teams to translate requirements into scalable data solutions; create documentation and handover materials.

Details

CompanyTata Consultancy Services
LocationHyderabad g.p.o., India
TypeFULL TIME
Nichetech
Experiencepermanent

Similar Jobs

P

Clinical Data Management Associate-Freshers | Bengaluru / Hyderabad / Chennai

Proxima Skills Academy

R

SAP Vistex Functional Consultant

Recmatrix Consulting

W

DFT Engineer

Wenger & Watson

H

Sales Engineer

Hakke Industries

T

Senior QA Engineer [T500-29096]

Talent500