AI Research

TRIBE v2: Meta's Predictive Foundation Model That Predicts Brain Responses to Sight, Sound, and Language

Meta FAIR's TRIBE v2 is a trimodal foundation model trained on fMRI data that predicts high-resolution human brain activity in response to video, audio, and language — enabling zero-shot predictions for new subjects and tasks.

A
AIDeveloper44 Team
May 16, 2026·5 min read
TRIBE v2: Meta's Predictive Foundation Model That Predicts Brain Responses to Sight, Sound, and Language

What Is TRIBE v2?

Meta FAIR and Jean-Rémi King's Brain & AI team have introduced TRIBE v2 (Trimodal Brain Encoder v2) — a foundation model trained to predict how the human brain responds to almost any sight or sound. Built on large-scale fMRI data, TRIBE v2 can predict high-resolution brain activity across vision, audition, and language, enabling zero-shot predictions for new subjects, new languages, and entirely new tasks it wasn't explicitly trained on.

Why a Foundation Model for the Brain?

Neuroscience has traditionally required months of lab work to map how a single person's brain responds to specific stimuli. Every new subject, language, or task required expensive new data collection. TRIBE v2 changes this by acting as a general-purpose brain encoder — a digital mirror of human brain activity that generalizes across individuals and modalities.

The model delivers fMRI predictions at 70x higher resolution than previous approaches, and the zero-shot generalization means researchers can make predictions about new populations and scenarios without collecting new neural data. What used to take months of lab work now takes seconds.

The Trimodal Architecture

TRIBE v2 is trimodal — it processes video, audio, and language simultaneously and maps all three to predicted brain activity patterns. This reflects the reality of how the brain processes the world: not as separate channels, but as an integrated multimodal experience. The model's ability to handle all three modalities jointly makes it uniquely powerful for studying cross-modal brain responses — for example, how the brain integrates spoken words with facial expressions in a video.

Research and Practical Applications

TRIBE v2 opens new research directions in cognitive neuroscience, BCI (brain-computer interface) development, and AI interpretability. By understanding which features of AI models most closely match brain representations, researchers can build more cognitively aligned AI systems. The model and research paper are available through Meta AI's research page, and an interactive demo is live at aidemos.atmeta.com/tribev2.

Enjoyed this?

Get more posts like this delivered to your inbox.

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode