Scan your resume against ATS criteria for this Lead Data Engineer role at UNI1066.
The U.S. Pharmacopeial Convention (USP) USP is an independent scientific organization that collaborates with the world's top experts in health and science to develop quality standards for medicines, dietary supplements, and food ingredients. USP's fundamental belief that
manifests in our core value of
through our more than 1,100 talented professionals across five global locations to deliver the mission to strengthen the supply of safe, quality medicines and supplements worldwide.
The Digital product engineering team at USP is seeking a Data Engineer with expertise in building robust data pipelines, managing large-scale data infrastructure, and enabling advanced analytics capabilities. This role is critical to supporting projects that align with our mission to protect patient safety and improve the health of people around the world. We are looking for a professional who understands the value of clean, accessible, and well-organized data, and who enjoys designing systems that empower data scientists, analysts, and business stakeholders to derive actionable insights. The ideal candidate will be passionate about data architecture, cloud technologies, and scalable engineering solutions that drive innovation and impact.
Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs.
Implement and maintain data governance frameworks and data quality checks within data pipelines to ensure compliance and reliability.
Build and optimize data models and data marts to support self-service analytics and reporting tools such as Tableau, Looker, and Power BI
Partner with data scientists to operationalize models by integrating them into production-grade pipelines, ensuring scalability, performance, and maintainability.
Collaborate with cross-functional stakeholders (business, product, analytics, and engineering teams) to translate business requirements into scalable data solutions and prioritize data initiatives.
Provide technical leadership and architectural guidance for data platform design, ensuring alignment with enterprise standards and long-term data strategy.
Lead design reviews, code reviews, and data architecture discussions to ensure best practices, reusability, and high-quality deliverables.
Collaborate with platform, DevOps, and security teams to ensure secure, cost-effective, and scalable data infrastructure.
Influence data roadmap and strategy by identifying opportunities for data platform enhancement, automation, and cost optimization.
The successful candidate will have a demonstrated understanding of our mission, commitment to excellence through inclusive and equitable behaviors and practices, ability to quickly build credibility with stakeholders, along with the following competencies and experience:
Education
Bachelor’s degree in relevant field (e.g. Engineering, Analytics or Data Science, Computer Science, Statistics) or equivalent experience.
Experience
7+ years of experience in big data technologies such as Python, PySpark, and SQL for processing structured, semi-structured, and unstructured data.
Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, Postgres, Neo4j (Azure/GCP equivalents such as ADLS, Synapse, ADF acceptable).
Experience in building batch, micro-batch, and streaming pipelines (real-time / near real-time) using Lambda/Kappa architectures.
Hands-on expertise in designing and delivering enterprise-scale data platforms, including data lakehouse, data warehouse, data lake, and data marts.
Strong understanding and hands-on implementation of data modeling techniques including - Data Vault 2.0, Dimensional Modeling, Knowledge Graphs, Big Table (OBT) approaches (Certification in at least area preferred)
Experience with medallion architecture and metadata-driven data pipeline frameworks.
Strong expertise in data governance frameworks, including- Data discovery, Data quality, Data security, Hands-on experience with DQ tools such as Great Expectations, Pydantic, etc.
Strong SQL and programming skills for data transformation, modeling, and analysis.
Hands-on experience in building and maintaining complex ETL/ELT pipelines and managing day-to-day data operations.
Experience with workflow orchestration tools such as Airflow (2+ years or equivalent).
Knowledge of streaming/event-driven architectures and modern data processing patterns.
Good understanding of dashboarding and visualization techniques with tools like Tableau, Power BI, or equivalent.
Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms.
Exposure to AI/ML concepts, with familiarity in Generative AI patterns (e.g., RAG, chunking techniques) as an added advantage.
Experience with scientific chemistry nomenclature or prior work experience in life sciences, chemistry, or hard sciences or degree in sciences
Experience with pharmaceutical datasets and nomenclature
Experience in developing Machine learning & Deep learning models; Familiarity with
and deploying ML models in production environments.
No
USP provides you with the benefits you need to protect yourself and your family today and tomorrow. From company-paid time off, comprehensive healthcare options to retirement savings, you can have peace of mind that your personal and financial wellbeing is protected
How to apply for Lead Data Engineer at UNI1066?
Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.
What is the salary for this role?
Salary details will be discussed during the interview.
What experience is required?
7+ years of experience is required.
Is this position still open?
Yes, currently active and accepting applications.
Explore related active roles in data_engineering
Lead Data Engineer
UNI1066 · Hyderabad