Russian scientists develop AI that reads emotions from voice and face
Scientists from the Higher School of Economics (HSE) in Moscow, together with researchers from the St. Petersburg Federal Research Center of the Russian Academy of Sciences and Sber’s Applied AI Center, have developed a neural network capable of analyzing multiple aspects of human behavior simultaneously.
The system can recognize emotional states, assess perceived personality traits and identify hesitation — situations in which a person's behavioral signals may indicate uncertainty or conflicting reactions.
Human communication extends far beyond words. Hesitation in the voice, changes in facial expressions, eye direction and even body posture can all add meaning to what a person says. The researchers are therefore training AI to analyze these signals collectively rather than treating each one separately.
AI Goes Beyond Facial Recognition
Many emotion-recognition systems focus primarily on facial expressions or speech. The new model takes a broader approach by processing multiple types of information, or modalities, simultaneously.
These include video, a person's voice, the content of their speech, as well as descriptions generated from facial expressions, eye movements, body posture and gestures. By combining these signals, the researchers aim to make AI analysis more closely reflect how people communicate in real-world situations.
For example, someone might say, “Yes, I’m completely comfortable with that,” while hesitating, avoiding eye contact and speaking in an uncertain tone. In such a situation, words alone provide only part of the available information. A multimodal AI system can examine several of these behavioral signals together.
However, the system does not actually read a person's thoughts. Instead, it analyzes observable patterns and generates estimates based on patterns learned from training data.
This distinction is particularly important when assessing personality. The researchers describe the system as evaluating perceived personality traits — how a person's behavior may appear based on available signals — rather than determining their true underlying personality.
Why Combining Different Signals Is Difficult
Developing such a system presented an unusual technical challenge. The researchers did not have access to a single large dataset in which all subjects had been labeled for emotions, personality traits and behavioral inconsistencies.
Instead, they worked with several datasets, or corpora, originally created for different tasks. Each dataset contained its own type of labels, while different information sources were not equally useful for every task.
According to project leader Elena Ryumina, video is particularly useful for estimating perceived personality traits, while the actual content of speech becomes more important when identifying inconsistencies.
The researchers therefore designed the neural network with both shared and task-specific components. The shared components help the system learn relationships between different types of information, while specialized components focus on the signals most relevant to a particular task.
This approach allows knowledge acquired from one dataset to be transferred to another task, rather than requiring the AI to learn each problem independently.
Learning Multiple Tasks at Once
The model was trained using a supercomputer at HSE. During training, the researchers combined several differently labeled datasets while teaching the system to perform three behavioral-analysis tasks.
The goal was not simply to achieve high performance on examples the neural network had already seen. The researchers also wanted the model to generalize its knowledge to previously unseen data.
According to the research team, integrating information from multiple datasets improved the system's ability to recognize emotions when processing unfamiliar data.
Such generalization is essential for practical AI systems because human voices, facial expressions, speaking styles and gestures vary widely. A system that works only with people or situations resembling its training data would have limited real-world usefulness.
Potential Applications
The researchers see potential applications in digital education, intelligent interfaces and user-support technologies.
In an educational setting, for example, future AI tools could take behavioral signals of confusion, hesitation or low emotional engagement into account rather than relying solely on whether a learner's answer is correct or incorrect.
Similarly, intelligent interfaces could adapt their interactions with users based on the behavioral information available to them.
The research team is now working toward a more flexible architecture. Their next step is to develop specialized neural-network “experts” capable of assigning different weights to voice, speech content, facial expressions, gestures and other signals depending on the task.
In parallel, the researchers are developing personalized digital avatars that could adapt their facial expressions, speech and behavior to a user's emotional state. Instead of responding only to the words people type or say, future interfaces could become more sensitive to how those messages are communicated.
The major challenge, however, is ensuring that such systems can interpret behavioral signals accurately and reliably. Human emotions and behavior are highly complex, and the ultimate value of this technology will depend not only on how much information AI can detect, but also on how reliably it can interpret and analyze that information.