Building a Scalable Data Strategy for CMC Development Sciences

Sakthi Prasad T, Content Director
Sep 25, 2026

Share This Post

Speaker: Christian Pilger, Director, Solution Architecture Lead, AbbVie

Zifo’s SiEE event 2026, Basel

Christian shares how his team is creating a scalable data foundation for CMC development sciences. Their goal is to turn fragmented scientific information into structured, standardized, and reusable data.

Development science transforms promising molecule candidates from discovery into commercial medicines. Along the way, information is generated across ELNs (electronic lab notebooks), laboratory systems, SAP, custom tools, instrument files, and spreadsheets. Christian explains that the greatest challenge is not data volume or velocity, but variety. Data is stored in silos, systems describe the same concepts differently, and inconsistent terminology makes information difficult to combine and analyze.

To address this, Christian’s team follows a “three gears” approach. They begin with scientific architecture by identifying and prioritizing specific business use cases. Next, they assess the current data landscape, define requirements, and improve the related processes. Only then do they design the systems architecture.

This sequence keeps the work focused on genuine scientific needs rather than technology. It also helps the team deliver tangible outcomes quickly while building solutions that can scale.

One early use case involved small-molecule formulation data, which scientists had traditionally recorded in Excel. The team introduced a structured ELN template and built an ETL pipeline to transfer the information into a data warehouse. The solution was later expanded from small molecule to biomolecule formulations and from basic composition data to complete recipes that included process details. As Christian notes, understanding a formulation requires knowing not only what it contains, but also how it was produced.

The team applied a similar approach to analytical requests. Previously, requests arrived through paper, email, handwritten notes, and labels on laboratory containers. An electronic request system created a more consistent experience for requesters and analytical teams. It was later extended to support stability studies by automatically generating requests when stored samples needed to be retrieved and analyzed.

Batch genealogy provided another opportunity. Christian’s team aligned SAP data with its scientific models and created a graph-based structure to represent relationships between batches. Beyond visualization, this graph foundation could support composition calculations and future AI applications.

Standardization is central to the architecture. The team uses data contracts to define common metadata for materials, batches, samples, containers, process steps, tests, and results. These contracts help align equivalent concepts across systems and create consistency around essential information such as identifiers, what was created, who created it, and when it was created.

Terminology presents a more human challenge. Different teams may use “clarity,” “turbidity,” or “opacity” for the same measurement. Instead of forcing everyone to adopt a single term, Christian’s team maps familiar local names to controlled parameters. Scientists can continue using familiar local terminology, while related information is connected through a shared parameter ID.

These parameters are organized into a taxonomy connecting the scientific question, the experimental test, and the parameter that provides the answer. This structure makes scientific data easier to find, compare, and reuse.

Christian emphasizes that technology is only part of the solution. Moving scientists away from spreadsheets, governing terminology, preventing duplication, and integrating independently designed systems require careful change management.

The approach is deliberately pragmatic: begin with a real scientific problem, standardize what matters, preserve familiar language where possible, and scale proven solutions gradually. The result is a flexible, platform-independent data foundation that can support visualization and advanced analytics while preparing the organization for future AI applications.

BOSTON
APR 13, 2027
Boston, MA
REGISTRATIONNOW OPEN
BASEL
JUN 1, 2027
Basel, Switzerland
REGISTRATIONNOW OPEN