CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Data Engineer - Pharma R&D
Roche
Roche

Data Engineer - Pharma R&D

Hyderabad, Telangāna
5-8 years exp
Full-time
Posted 5d ago
2 views
Actively Hiring Urgent Opening Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Data Engineer - Pharma R&D role at Roche.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

  • Job Description:

The Data Engineer – Clinical Study Design sits at the intersection of data architecture, clinical science, and technology delivery, helping build the data foundation behind Study Designer, a digital product transforming how studies are designed. This role combines strong data engineering skills, an understanding of clinical study data and workflows, and technical curiosity to translate complex clinical data structures into reliable, scalable pipelines and models that power the platform's insights. The ideal candidate is curious, collaborative, AI-minded, and passionate about using data to enable smarter, faster, and more effective clinical study design.

Key Responsibilities

Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization

Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships

Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population

Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores

Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)

Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer

Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations

  • Preferred Qualifications:

Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.

5-8 years of experience building production grade data platforms and pipelines.

Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)

Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data

Apache Spark or Databricks for batch processing of large document

Required Skills

Python — Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)

SQL — Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes

Graph Databases — Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language;/knowledge graph modeling for biomedical entities

AWS — Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration

NLP / Document Processing — Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, pgvector, or Pinecone)

FastAPI — Building data serving endpoints; async patterns; integration with the application backend

AI/ML Data Infrastructure — Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns

Pipeline Orchestration — Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry logic, and monitoring

CI/CD & IaC — Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS

#Hyderabad2026

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.

Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer.

Key Requirements & Skills

  • Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.
  • 5-8 years of experience building production grade data platforms and pipelines.
  • Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)
  • Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data
  • Apache Spark or Databricks for batch processing of large document
  • Python — Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)
  • SQL — Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes
  • Graph Databases — Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language;/knowledge graph modeling for biomedical entities
  • AWS — Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration
  • NLP / Document Processing — Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, p
  • FastAPI — Building data serving endpoints; async patterns; integration with the application backend
  • AI/ML Data Infrastructure — Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns
  • Pipeline Orchestration — Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry l
  • CI/CD & IaC — Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS

Benefits & Perks

medical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)**

Frequently Asked Questions

How to apply for Data Engineer - Pharma R&D at Roche?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

5-8 years of experience is required.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
Roche

Roche

About | University of Rochester Skip to content Search Close Menu Close Academics Research Admissions Campus Life Medicine About Apply Visit Give Careers Search Discover the University of Rochester tag" alt="photo of Rush Rhees Library and Eastman Quadrangle" /> Making the world ever better As a world-leading research institution, the University of Rochester brings creative problem-solvers together across disciplines to pursue bold questions that lead to big impact. We advance research. We pursue lifesaving cures. And we accelerate innovations that make our lives—and the future—ever better

Visit Company Website

More jobs at Roche

Regulatory AI & Automation Specialist

Hyderabad

Business Operations Associate (Marketing)

Hyderabad

Business Operations Associate (Marketing) (Arabic speaking)

Hyderabad

Share this Opening

Job Alerts for data_engineering

Receive email alerts whenever new data_engineering roles in Hyderabad are posted.

Set Free Alert →

Similar Openings

Explore related active roles in data_engineering

View all
UrgentActively Hiring
Momentum Financial Services Group
Lead Data Engineer
Momentum Financial Services Group Verified
10+ years
Salary not disclosed
Hyderabad (Remote)
data_engineeringFull-timeRemote
Posted 1d ago
Apply Now
UrgentActively Hiring
aecom2
Data Engineer
aecom2 Verified
0-2 Yrs
₹111/mo
Bristol, 3 RIVERGATE, gb
data_engineeringFull-time
Posted 1d ago
Apply Now
UrgentActively Hiring
Magnals
Lead Data Engineer
Magnals Verified
5+ years
₹10.7L – ₹12.1L/mo
Remote
data_engineeringFull-timeRemote
Posted 1d ago
Apply Now

Data Engineer - Pharma R&D

Roche · Hyderabad

Apply on Company Website