Job Description – Data Engineer
Role Overview
We are looking for an experienced
Data Engineer
with 5–7 years of hands-on experience in
ETL/ELT, database engineering, Big Data processing, and complex bi-directional data integrations
. The candidate should have strong SQL and database expertise, experience working with large datasets, and the ability to design, develop, optimize, and support reliable data pipelines in production environments.
Key Responsibilities
- Design, develop, and maintain scalable
ETL/ELT pipelines
for batch and near-real-time data processing.
- Build data integrations using
REST APIs, SFTP, files, webhooks, and database-to-database interfaces
.
bi-directional data integrations
between applications, databases, and external enterprise systems.
- Implement data synchronization, incremental processing, reconciliation, error handling, retries, and duplicate/missing-data prevention.
- Develop complex
SQL queries, stored procedures, functions, views, and data transformations
.
PostgreSQL and SQL Server
; exposure to MongoDB is preferred.
database and query performance tuning
, including execution-plan analysis, indexing, joins, partitioning, locking/blocking, and long-running query optimization.
- Work with large datasets and distributed processing frameworks; optimize data processing for scalability and performance.
- Implement data validation, quality checks, logging, auditing, and monitoring for critical pipelines.
- Troubleshoot production ETL, database, integration, and data-quality issues and perform root-cause analysis.
- Collaborate with application, product, DBA, DevOps, and client teams to resolve end-to-end data issues.
Required Technical Skills
- 5–7 years of Data Engineering / ETL experience
- Strong
SQL and Python
skills
- Strong hands-on experience with
PostgreSQL and SQL Server
ETL/ELT and data pipeline architecture
complex data integrations and data synchronization
- Experience with REST APIs, SFTP, file-based and database integrations
- Strong understanding of
database performance tuning and optimization
- Experience handling large-volume datasets
- Good understanding of data modeling, data quality, validation, and reconciliation
- Strong troubleshooting and problem-solving skills
Apache Spark / PySpark
Apache Airflow
Talend / SSIS / other ETL tools
MongoDB
- Azure / AWS / GCP
- Azure Databricks / Delta Lake / Data Lake
- Kafka or other streaming technologies
- CDC / incremental data processing
- Git, Azure DevOps/GitHub and CI/CD
Preferred Candidate Profile
The ideal candidate should be comfortable working across the complete data lifecycle:
Candidates who have experience working with
multiple databases, high-volume data, complex bi-directional integrations, and performance-critical production workloads
will be preferred.
Key Competencies
- Strong analytical and problem-solving ability
- Good understanding of data architecture and integration patterns
- Ability to independently troubleshoot complex data issues
- Ownership of deliverables and production issues
- Good communication and cross-functional collaboration skills
- Education: Bachelor's/Master's degree in Computer Science, IT, Engineering, or related technical discipline.