菠萝视频

>

Building the World’s Largest Biomedical Informatics Enterprise

By Paul Govern
Biomedical informatics illustration
ILLUSTRATIONS BY HANK OSUNA

 

One of the many things a person can do with a 菠萝视频 University network ID and password is explore the Record Counter, or RC. Faculty, staff and students interested in medical records can ask the RC just about anything they鈥檇 like鈥攁s I have.

In a research database containing some 2 million de-identified patient records from (around 18 years鈥 worth), there are, for example, 268 records indicating assault by human bite鈥136 females and 132 males. That was the count as of May 18, 2014.

I first noticed this tersely evocative patient label 13 months earlier, and it turns out the RC鈥檚 count has since risen, unhappily, by 22.

菠萝视频ers use the RC to check whether the database contains cohorts of sufficient size to probe a wide range of biomedical questions. To step any further into this secure, otherwise restricted data warehouse, as it鈥檚 called, investigators sign data-use agreements and obtain approval to pursue specific hypotheses.

菠萝视频 claims the world鈥檚 largest biomedical informatics enterprise. With great purpose and deftness, the university has turned its EHR (electronic health record) into an object of study.

Kevin Johnson ID card鈥淭he beauty of data analytics is that we can run initial experiments in data without involving any patients at all,鈥 says , Cornelius 菠萝视频 Professor of Biomedical Informatics, chair of the and professor of pediatrics.

鈥淭hat鈥檚 the really seminal observation that has been made by many industries before health care even got involved: that behaviors and patterns exist that can be predictable even without fully understanding why they occur, and if you can model those patterns, you can effect change.鈥

Most Americans have been familiar with the concept of data mining鈥攖he task of generating new information from large databases鈥攁s the enterprise behind Web searches or as a stratagem for predicting consumer behavior. In the medical research domain, by contrast, the pickings have been comparatively slight, with technology of the sort that powers commerce and banking missing from hospital rooms and clinics.

But no longer: The American Recovery and Reinvestment Act of 2009 included $19 billion for health information technology. This largesse, together with decreasing costs for DNA sequencing technology, spells employment for a certain type of data expert, and unprecedented opportunities for research.

鈥淪ome studies suggest that we have rigorous evidence for only 15 to 20 percent of what we do in the hospital and clinic. The rest is opinion or extrapolation from other data,鈥 says , associate professor of medicine and director of the .

Russell Rothman ID cardBiomedical informatics is a science that draws connections between data and medicine, whether those data concern diseases, health care processes or human biology in the form of genomics and proteomics. Everyone who studies health records has the same goal: more precise medicine, leading to improved patient outcomes.

VUMC was a relatively early investor, adopting routine electronic record keeping in 1995. The Record Counter sits atop the Synthetic Derivative, where records are stripped of personal identifiers and, without sacrificing their scientific utility, are randomly altered to help prevent re-identification of specific patients. Under a program called BioVU, launched in 2007, de-identified records are linked to de-identified DNA samples representing 180,000 patients and counting.

At first, using electronic medical records to study disease can look like a hopelessly vexed proposition. Hospital and clinic documents are geared toward serving patients, care teams and billing departments. In relation to any disease states that may be present in a population, the data in medical records can be noisy and sparse. Data that spring up in the course of patient care may droop and fade under the lens of post-hoc data science. Clinical lab tests or other measurements that may be of interest seem to occur only irregularly, if at all. Diseases frequently sprawl atop one another鈥攚hich is interesting, but renders everything less computationally distinct. Interpreting the occurrence of drugs is apt to call for some suspension of disbelief. Text, where some of the richest information is said to reside, requires sophisticated processing to yield computable terms, and the same goes for medical images.

Just as field biologists are prepared to contend with crocodiles and mosquitoes, data scientists who study medical records are prepared to overcome riotous uncertainty. 菠萝视频 investigators are devising new methods for using medical records鈥攐r health records, as doctors are coming to call them鈥攊n rigorous analyses of health and disease.

Josh Denny ID card鈥淎s time goes on, the number of things you can analyze with genotype sets and electronic medical-record populations is limited only by the reasons people see their doctors,鈥 says , BS鈥98, MD鈥03, MS鈥07, associate professor of biomedical informatics and medicine and director of the . 鈥淭he clinical data can be messy, but contain perhaps the single richest source of disease history, drug exposures and their response, and prognosis available for research.鈥

While VUMC has long been recognized as a hub for innovative data science, 鈥渢he big difference during the past few years,鈥 Johnson says, 鈥渉as been that the computational tools are much more available, computational expertise is much more available, and computers in general are more capable, which has made some formerly impossible problems much more tractable.鈥

UNSUPERVISED, SEMI-SUPERVISED AND SUPERVISED LEARNING

All sorts of useful upstream signals of downstream risk are there for the finding in the EHR. In broad terms, one approach to these data is to set aside momentarily the clinical labels ascribed to patients in the hospital and clinic, to see where patients amass with regard to more stripped-down, straightforward information like clinical lab results and exposure to medications.

It鈥檚 a bit like looking through the wrong end of the telescope, but it鈥檚 human biology nonetheless. It involves circling back to view two graphs of the population, labeled and unlabeled. When these graphs are superimposed, a familiar label like, say, rheumatoid arthritis, might be seen to break up into population clumps, or may overlap with other diseases. Call that newly revealed structure a demonstration of unsupervised learning. The clumps and overlaps, interrogated in the lab, may or may not yield new biological or epidemiological stories.

Another approach鈥攃all it supervised learning鈥攎ight start by using the EHR to infer clinical labels as carefully as possible, giving consideration to the full record and employing natural language processing to extract concepts from any text therein. Finding cases and controls in an EHR population鈥攔heumatic and non-rheumatic, for example鈥攃an serve as a starting point for exploring the role of genetic variation or studying therapies. Unsupervised, semi-supervised and supervised learning work in concert.

Dan Byrne ID card鈥淢achine learning, unsupervised learning and artificial intelligence sound sexy but are no match for real intelligence with clinical input and real-time predictive models that have withstood rigorous validation,鈥 contends , a senior biostatistician at the .

Under a program called Cornelius, Byrne and colleagues are testing the usefulness of EHR-based patient risk stratification. Randomized controlled trials using predictive models of pressure ulcers (bedsores) and hospital readmission (within 30 days of discharge) are underway in 菠萝视频 University Hospital. In the pipeline are models of urinary tract infection, embolisms (blood clots), bloodstream infection and patient falls.

鈥淲e鈥檙e building the infrastructure to help 菠萝视频 investigators move beyond simply publishing a paper about a predictive model to using it to improve outcomes in a sustainable way,鈥 Byrne says.

, assistant professor of biomedical informatics, arrived at 菠萝视频 from Google in 2010. When a medical symptom is entered using Google, it fires search technology conceived and initially developed by Lasko.

鈥淚 thought 菠萝视频 was the best place in the world for this kind of research, and I still think it鈥檚 one of the best,鈥 he says.

Tom Lasko ID card鈥淪upervised learning is a great technique, but it only looks where you tell it to look, and you鈥檙e limited by your preconceived notions of what causes a given disease,鈥 observes Lasko, who has led a demonstration of so-called 鈥渄eep learning鈥 on EHR data. It鈥檚 an unsupervised learning technique inspired by visual processing in the brain.

Lasko鈥檚 search for precision happens to begin far out in the land of hidden phenotypes. (If eye color is a trait, for example, then blue, brown or hazel eyes are phenotypes. Any feature or pattern in an organism might qualify as a phenotype, including states of health and disease.)

Lasko gathered records from 4,368 patients, half with gout and half with leukemia. Both types of patients experience elevated uric acid levels and receive repeated testing. To enable unsupervised learning on just these data, Lasko computed longitudinal probability distributions for each patient鈥檚 uric acid levels鈥攖hat is, he transformed noisy, irregularly timed uric acid snapshots into continuous graphs that are more suggestive of underlying disease processes.

He processed this unlabeled information with a deep-learning algorithm, took the resulting population features as new inputs, processed those, and arrived at a Rorschach-like graph of the population. Then he retrieved the disease labels he had set aside and used them to color in the graph.

In this final picture, not only do gout and leukemia break cleanly apart鈥攑icture Lasko as Charlton Heston standing resolutely before Cecil B. DeMille鈥檚 Red Sea鈥攂ut the two labels also break up into substructures. And the point of this demonstration is that these sub-groupings might carry meanings of their own.

鈥淭hat I鈥檓 finding this clump, or this area in the data space where people seem to be congregating, is not a proof of anything,鈥 Lasko explains, 鈥渂ut it鈥檚 a strong indication that something mechanistically common may be underlying that.

鈥淵ou could hand this information to a geneticist or some other researcher. Or maybe this is an opportunity for a clinical trial,鈥 he adds. 鈥淚f we see a clump responding great to a particular drug for this disease, then maybe we should go straight to testing whether the effect is real and whether it could be used in clinical medicine.

鈥淢y point is that precise data-driven definitions of what a disease is are more likely to be correlated with the underlying pathophysiologic mechanism than our clinically driven definitions. I haven鈥檛 proven that yet鈥攂ut that鈥檚 what I鈥檓 going after.

鈥淢y ideal setup,鈥 he adds, 鈥渨ould be to have everybody鈥檚 information in the world.鈥

REDEFINING PHENOTYPES

When patients respond differently to the same drug for the same diagnosis, as happens so often, distinct phenotypes might be in play. A subtext of unsupervised learning on EHR data is that many more diseases may exist than are contemplated in current medicine.

Brad Malin ID cardIn a study now at press, and colleagues map an EHR population with reference to links they鈥檝e managed to establish between individual drug prescriptions and their precipitating diagnoses. All clinical phenotypes are in play in this demonstration, and the one under examination is hypertension.

If this type of approach were to produce patient clusters that simply match known phenotype labels, then 鈥渇antastic,鈥 Malin says, 鈥渂ut our expectation is that there鈥檚 much more complexity to these patients. We are investigating the nuance of what makes patients different so that we can redefine the phenotypes.鈥

Malin, associate professor of biomedical informatics and computer science, directs 菠萝视频鈥檚 . 菠萝视频 on the EHR is a newish pursuit and, for Malin, the methodology itself is the story鈥攊t鈥檚 where the important novelty lies. But studies highlighting new informatics methodology are largely relegated to journals read only by other data scientists. Breaking through requires demonstrating a method in a population. 鈥淭hat鈥檚 one of the things 菠萝视频 is capable of doing that a lot of other places are not,鈥 he says.

A phenome is the sum of phenotypes to be observed in an individual or species. In a phenome-wide genetic association study, or PheWAS, appearing last year in Nature Biotechnology, Josh Denny and colleagues borrowed repurposed genotype data from 13,835 patients from five different medical centers around the country.

Continue reading related stories.

Informatics for the Classroom and the Operating Room

Images to Algorithms

According to billing codes, these patients collectively exhibited 1,358 different diseases and conditions鈥攖hat is, they represented the breadth of the clinical phenome. Denny focused on 3,144 common genetic variants already implicated in one disease or another, measuring the frequency of each variant in each disease group, comparing them to the general population.

The study replicated many known gene鈥揹isease associations. The real payoff, however, was the discovery of 63 previously unknown ones, each an example of a genetic variant having independent association with more than one trait鈥攑leiotropy, as it鈥檚 called.

The New York Times covered the study as the first large-scale PheWAS.

鈥淚f you want to say this is a coming-out party for PheWAS, then it鈥檚 also in some ways a coming-out party for the electronic medical record as a tool for genetic studies,鈥 says Denny, who had first demonstrated the feasibility of such a scan in 2010.

Denny and others at 菠萝视频 also pursue the more familiar inverse approach of identifying subjects with and without a given disease, scanning the breadth of their genomes, and checking the frequency of genetic variants against the general population.

A 2011 study of hypothyroidism by Denny and colleagues was the first GWAS (genome-wide association study) of a disease using the EHR and repurposed genotype data from previous scans. 鈥淥ur premise was, let鈥檚 see if we can basically do a 鈥榥o genotyping鈥 GWAS,鈥 says Denny. 鈥淐an we use what鈥檚 already on the shelf, pick another disease, and analyze it within those samples?鈥

Near a gene that codes for a thyroid transcription factor, they identified four common genetic variants as being highly associated with primary hypothyroidism.

These super-efficient approaches to discovery form the rationale behind the eMERGE (electronic medical records and genomics) Network, a national consortium of biorepositories linking DNA samples to de-identified medical records. 菠萝视频 is the network鈥檚 coordinating center, and BioVU is by far the largest repository. According to a recent study, the median cost of BioVU studies is less than one-17th that of similar studies performed elsewhere, and while BioVU studies take a median time of three months to identify subjects, the median grant period for similar studies is three years.

The provisos attached to the $19 billion in federal incentives for health information technology include adding to the EHR more structured, machine-readable information about the clinical process. Meanwhile, 鈥渢he only place you鈥檙e going to get the history of what brought patients to you, leading to a given diagnosis, is natural language processing,鈥 says Denny.

Natural language processing (NLP) uses a combination of linguistic rules and statistics. In an example of machine learning, the computer digests an exhaustively hand-annotated corpus, yielding a statistical model of written English as used, in this case, in physician notes, nurse notes, messages from patients and so on. 鈥淚f you want to look at what someone鈥檚 first presenting symptom was for multiple sclerosis, for example, that鈥檚 an NLP task,鈥 says Denny. 鈥淚f you want to know even when they were diagnosed, many times that鈥檚 also probably an NLP task.鈥

While commercial interest in clinical data analytics is bustling, some well-known companies have had only limited success moving new health-record technology.

Trent Rosenbloom ID card鈥淭he largest software companies are not making much headway into the electronic health-record space,鈥 says , MD鈥96, MPH鈥01, whose research includes evaluation of health information technology. 鈥淕oogle shied away, Apple shied away, Microsoft has had limited usage.

鈥淚 suspect that some new startup that gets the right investor will disrupt the field,鈥 Rosenbloom says. 鈥淢y impression of where we鈥檙e going is that ultimately EHRs and related applications will become these very small, App Store-like things that live on your phone and do things that are highly individualized, and the data will all be fairly standardized and live elsewhere, in the cloud.鈥

It鈥檚 a far cry from where things stood back in 1998, when 菠萝视频鈥檚 rounded collection of experts in this field could fit around dining room table.

Bill Stead IDStead is the McKesson Foundation Professor of Biomedical Informatics, associate vice chancellor for health affairs, and VUMC鈥檚 chief strategy and information officer. The fact that biomedical informatics is artfully sewn into both the clinical enterprise and the biomedical research enterprise is the consequence of a vision conceived and fostered by Stead these past 23 years.

鈥淭his stream of research gives us hypotheses, feature extraction and phenotype signatures,鈥 says Stead. 鈥淣ow let鈥檚 put that together. Let鈥檚 create a discovery platform that provides a gold standard to help us extract and export executable knowledge that can be incorporated into everybody鈥檚 electronic health systems and life management applications.

鈥淭he end-game vision,鈥 he concludes, 鈥渋s executable knowledge to support 鈥榳hole person鈥 health and health care.鈥


Paul Govern, an information officer at 菠萝视频, writes about the VUMC clinical enterprise, including efforts to improve the quality, cost and safety of health care delivery. He also helps cover medical research, bioinformatics and clinical informatics at the Medical Center.


Watch a presentation by Bill Stead about electronic health records as a platform for research: