Section head: Prof. dr. M.A. van de Wiel (mark.vdwiel@amsterdamumc.nl)
In biomedical studies, we often encounter a special form of Big Data: one in which we have more measurement points (variables, such as genes or proteins) than individuals (patients)—these are known as “omics” studies. This requires specialized statistical and machine learning techniques to, for example, build predictive models or test hypotheses. In addition, we are increasingly incorporating external information—such as data from databases—into our analyses. This helps us make our predictive models more accurate and robust. We also look at the bigger picture: how do the variables interact with one another (e.g., genes with proteins), whether one-on-one or within networks.
In addition, we use Big Data that is truly “big” in terms of the number of individuals involved. Here, we specialize in causal inference: what methods can we use to draw conclusions about cause and effect? We also use statistical techniques to link (anonymized) data sets together so that we can utilize even more data to answer our research questions.
We therefore develop methods to perform all these complex analyses and to test and validate the results.