In This Section
CMU Cognitive Psychologist’s Work Has Generated Over 14,000 Language Studies
TalkBank began as a project of convenience in the 1980s, but has grown into a family of databases that has transformed the field of language science.
By Jason Bittel Email Jason Bittel
- University Communications and Marketing, Media Relations
- Email media-relations@andrew.cmu.edu
Most people have never heard of TalkBank. But for those working in the language sciences, the family of 14 databases created by Carnegie Mellon University’s Brian MacWhinney has become the foundation for research into everything from how children learn to speak and how people pick up a second language to patterns that can reveal the onset of dementia.
At their heart, TalkBank and its spinoffs — including AphasiaBank and DementiaBank — are enormous, standardized databases of audio and video recordings of people talking that allow scientists to study many different components of language.
Before these databases, scientists relied on their own, often highly personalized transcription formats, codes and language analysis programs. And like the aftermath of the Tower of Babel, this severely limited how much cooperation researchers working in any given lab could offer one another.
But now, after nearly five decades, the TalkBank Project is the world’s largest open-access integrated repository for spoken-language data. It has ushered in a revolution to the language science field and yielded more than 14,000 scientific publications.
“We have people all over the world, with enormous amounts of data, and everybody’s using it,” said Brian MacWhinney, Teresa Heinz Professor of Cognitive Psychology in CMU’s Dietrich College of Humanities and Social Sciences and creator of TalkBank.
In fact, this work has been so transformative, MacWhinney was recently awarded the 2026 Yuen Ren Chao Prize in Language Science, a lifetime achievement award presented by the Faculty of Humanities of The Hong Kong Polytechnic University (PolyU).
The prize recognizes MacWhinney’s “lifetime of distinguished contributions to language science, encompassing integrative theoretical innovation, research infrastructure development and lasting international impact on the study of human language,” according to PolyU.
“Brian’s work is the gift that keeps on giving,” said Susanne Ferber, professor and head of the Department of Psychology. “His foundational contributions to science go well beyond understanding the complexities of first or second language acquisition to advance our insights into how spoken language offers a window into brain health.”
The Power of Standardization
MacWhinney still recalls the moment in 1981 when he and his early colleagues started down the path that he’d blaze for the next half-century.
“I can remember very clearly that we were at a meeting in Nijmegen, and we had some child language transcripts that had been mimeographed,” said MacWhinney, referencing the predecessor technology to the copier. “And we were marking them up with comments on the side, in pencil, and it occurred to me… they had just come out with the IBM PC. Why not make files that could be more quickly and easily circulated?”
By 1984, MacWhinney and his team had won the support of the MacArthur Foundation, which enabled them to bring 20 of the world’s top language researchers to Concord, Massachusetts, to hammer out the details of what would soon be known as the Child Language Data Exchange System, or CHILDES.
For the first time, CHILDES provided language researchers across the world with standardized data. But it wasn’t long before MacWhinney saw the need to build off of CHILDES’ success to develop still more databases that could be used for adults and other populations.
It was in 2001, with the support of a National Science Foundation Infrastructure Grant, that MacWhinney and his team created a standardized transcription and coding system, called CHAT, as well as a standardized analysis program, CLAN. Combined with the database, CHAT and CLAN made up the skeleton of TalkBank — a formula the team would repeat to create 14 more standardized databases that have become invaluable in the language sciences and which can be used across 18 languages.
“There’s a database for dementia, called DementiaBank. There’s a database for autism, called ASD Bank. There’s traumatic brain injury, there’s stuttering, there’s second language learning,” said MacWhinney.
Name a topic of interest in language science, and there’s now a standardized database for it.
What’s Next for Language Science?
As just one recent example of the advances being made in language science as a result of MacWhinney’s work, he pointed to a recent effort to develop better automatic speech recognition for children, who have distinct vocal characteristics, inconsistent pronunciation and who are still learning how to control their mouths when they speak.
Known as the On Top of Pasketti: Children’s Speech Recognition Challenge, over 800 language scientists submitted some 2,100 solutions to the problem prompts, which made use of data from the CHILDES and PhonBank databases in TalkBank. In the end, the winners blew MacWhinney away.
“The top solvers cut the error rate of the best existing children’s speech model by more than 40%,” he said. “That kind of increase in accuracy is incredible. In computer science, you might shoot for 1% or 2%.”
Similarly, MacWhinney believes there is enormous promise in the work related to speech patterns and dementia.
“It’s one area where we’ve only just started about five years ago, and it’s really the only publicly available set of data on language and dementia,” he said. “Several hundred computer science labs, as well as private companies, are now using it to develop methods for early detection.”
“The best part about it is you can do it without putting somebody in an MRI machine or extracting fluids from their spinal cord,” said MacWhinney. “You just take a measure of their speech, and you can really detect onset of dementia with high accuracy. It’s not perfect, but we’re talking up to 95% accuracy.”
Of course, TalkBank is also a natural partner for the fields of machine learning and artificial intelligence, including efforts to get AI models to learn language like children called BabyLLM.
“AI and machine learning need lots of data, and TalkBank provides it,” said MacWhinney. “At the same time, AI is reshaping TalkBank data, analysis and infrastructure.”
Even though updating and managing the databases now takes up so much of MacWhinney’s time that he can no longer do the experimental or theoretical work that got him into the field, he loves what he does.
“I’m not retiring, because it’s fun,” said MacWhinney, who is fluent in Spanish, Hungarian, German and French. “It’s fun to look at language.”