bebo Technologies is a leading complete software solution provider. bebo stands for 'be extension be offshore'. We are a business partner of QASource, inc. USA[www.QASource.com]. We offer outstanding services in the areas of software development, sustenance engineering, quality assurance and product support. bebo is dedicated to provide high-caliber offshore software services and solutions.
Our goal is to 'Deliver in time-every time'
Let's have a 360 tour of our bebo premises by clicking on below link:
Job Responsibilities:
- Design and implement scalable data platforms leveraging Data Lake, Lakehouse, Data Mesh, and modern data architecture patterns
- Build and optimize data pipelines for batch and real-time processing using Databricks, Apache Spark, dbt, and cloud-native AWS services
- Develop robust data ingestion frameworks for structured, semi-structured, and unstructured data from APIs, files, databases, and streaming sources
- Design and develop scalable Python-based microservices to enable secure and efficient data sharing across systems and applications
- Build RESTful APIs and event-driven services for exposing curated datasets from Data Lake and Lakehouse platforms
- Work extensively with Python, PySpark, and SQL for data transformation, processing, and data engineering workflows
- Implement streaming data pipelines using Kafka/Kinesis and integrate them with downstream analytics and data platforms
- Design and manage large-scale datasets using formats such as Parquet, JSON, CSV, and IoT/sensor data
- Optimize data storage, partitioning, and query performance for high-volume analytical workloads
- Design and implement data orchestration workflows using Apache Airflow and/or Databricks Workflows
- Implement infrastructure and data platform components using Terraform and Infrastructure as Code (IaC) practices
- Containerize data applications and services using Docker where applicable
- Collaborate with cross-functional teams, including Data Architects, Data Analysts, and BI teams, to operationalize Data Lake and Lakehouse solutions
- Contribute to data modeling in Lakehouse environments, including Medallion Architecture and dimensional modeling
- Implement data quality, validation, monitoring, and observability practices across data pipelines
- Use data quality and observability tools to identify data issues, monitor pipeline health, and improve data reliability
- Ensure data quality, reliability, security, and observability across data pipelines
- Implement serverless data processing and pipeline solutions where applicable
Job Requirements:
- 6–8 years of experience in Data Engineering or related roles
- Strong understanding of modern data architectures such as Data Lake, Lakehouse, Data Mesh, and Data Products
- Strong proficiency in Python, PySpark, and SQL
- Hands-on experience with Databricks, Delta Lake, and/or Snowflake
- Experience with ETL/ELT frameworks and data orchestration tools
- Hands-on experience with Apache Airflow and/or Databricks Workflows
- Practical experience with streaming technologies such as Kafka and/or AWS Kinesis
- Strong understanding of batch processing frameworks and technologies such as Apache Spark, AWS Glue, and dbt
- Proficiency in handling structured, semi-structured, and unstructured data
- Experience with modern data modeling techniques, including Star Schema, Snowflake Schema, and 3NF
- Familiarity with vector databases and data architectures supporting AI/ML use cases
- Hands-on experience with AWS services including S3, Glue, Glue Data Catalog, Athena, and Redshift
- Experience with Infrastructure as Code (IaC) tools such as Terraform
- Experience with containerization technologies such as Docker
- Experience with Git, CI/CD, and DevOps practices
- Experience implementing data quality, monitoring, and observability solutions using relevant tools and frameworks
- Good understanding of cloud security concepts, including IAM, encryption, access control, and data governance
- Strong understanding of Data Lakehouse concepts, architecture, and implementation patterns
- Familiarity with serverless architectures and Python-based serverless data processing solutions