Interhuman AIs coverbillede
Interhuman AI

Interhuman AI

Softwareudvikling

Social Intelligence for AI

Om os

Interhuman AI is the social intelligence layer for AI, enabling machines to understand the 93% of human communication that happens non-verbally. By analyzing facial expressions, body language, voice tonality, and conversational context in real-time, our technology transforms basic AI interactions into meaningful conversations. Our multimodal perception system works with any AI platform, enhancing interactions by detecting user confusion, engagement, and emotional states that text alone misses. This allows AI to adapt its approach mid-conversation, leading to more effective human-machine communication.

Websted
https://interhuman.ai
Branche
Softwareudvikling
Virksomhedsstørrelse
2-10 medarbejdere
Hovedkvarter
Copenhagen
Type
Privat
Grundlagt
2023

Beliggenheder

Medarbejdere hos Interhuman AI

Opdateringer

  • As AI moves from something we use occasionally to something we interact with every day, the way it understands us becomes increasingly important. We already communicate with AI through text and voice, but human communication has never been about words alone. We rely on tone, hesitation, expression, timing, and countless other signals to understand what someone actually means. If AI is going to become a more natural part of how we work, learn, create, and communicate, it will need to understand more than language. It will need to understand the signals that sit around it. That is the direction we are exploring with Inter-2: allowing AI products to get more insight from every conversation, allowing for a richer understanding of human behaviour and, ultimately, making human-AI interaction feel a little more like human interaction. Hear the conversation between Benjamin B. Biering, PhD, our Chief AI Officer and Alberto Cardaci, the Head of Behavioral Intelligence, on the matter at hand.

  • Building a model that understands social signals is not just a matter of scaling data or adding parameters. For Inter-2, our team took a different approach: we changed how the model learns. We designed a purpose-built training curriculum that progressively introduced the model to the problem, starting with output structure, then grounding signals in audio and visual information, and finally introducing increasingly complex social interactions. A few of the technical changes made a significant difference: - Counterfactual audio-visual training to distinguish what is supported by visual evidence from what can be inferred through sound, and ensure modality dropout robustness. - Multi-signal training to handle moments where several signals coexist, such as confidence and skepticism, or interest and disagreement. - Structured outputs that provide engagement status, detected signals, confidence levels, and explanations grounded in observable evidence. - Inference optimization that brings faster responses alongside higher accuracy. Inter-2 reaches 47.0 EW-F1 and 3.5× relative inference speed compared with our previous setup. The result is not simply a larger or faster model. It is a model trained specifically for a harder problem: extracting structured information about how people communicate from what they say, how they sound, and what can be observed. And this is only the surface. We wrote a deeper technical breakdown of the curriculum, post-training process, audio-visual grounding, and evaluation behind Inter-2. Read the full deep dive: https://lnkd.in/evYe4sdy

  • Today, we’re introducing Inter-2, our next step in socially intelligent AI. Inter-2 helps AI products understand how people communicate, beyond the words we use. It allows you to notice more by detecting social signals such as hesitation, agreement, confusion and skepticism in real time, with significantly higher accuracy, speed and robustness compared to Inter-1. For products built around human interaction, the model adds a layer of information that transcripts alone can’t capture allowing for a sales intelligence platform to detect how a customer is responding as a conversation unfolds; a user research platform to identify moments of hesitation or engagement; a coaching tool to understand not only the words being used, but how they are being delivered. Inter-2 is another step toward AI systems that can perceive and respond to the social dynamics of human interaction, helping products get more insight from every conversation and do more with it. Inter-2 is available today at interhuman.ai

  • Tomorrow, our Paula Petcu will join Aisel Health, for their conversation series, Minds & Mechanics to explore how technology can become a training tool, and how AI can move beyond language to better understand the people behind it. She’ll also share our work on FRED, developed together with researchers from Forskningsenheden ved Børne- og Ungdomspsykiatrisk Center, where AI is being explored as part of parent training and practical support for families. https://lnkd.in/eJgtS5kp

    Se organisationssiden for Aisel Health

    4.020 følgere

    Minds & Mechanics · This Friday · Technology as a training tool · Paula Petcu · Interhuman AI "AI understands language. It doesn't understand people." That's the gap Paula Petcu and Interhuman AI is building for. Paula is founder of Interhuman AI, the Danish startup building social intelligence infrastructure for AI - reading tone, hesitation, stress and body language in real time, not just words. One of Interhuman's use cases is in the intersection of mental health and technology: FRED - developed together with researchers from Forskningsenheden ved Børne- og Ungdomspsykiatrisk Center. FRED combines parent training, practical exercises and an AI coach grounded in research on emotion regulation and conflict management. The project is now being tested with families, an important step in understanding where AI can complement existing support. Curious to learn how that looks like? Join us for the Friday for Minds & Mechanics - Aisel's conversation series for the mental health ecosystem.

  • "You're nodding right now, for example. That is a signal that is important in this conversation. You're understanding what I'm saying, you're acknowledging, you're listening. All these signals we're trying to teach a model to pick up on." That is what Frederik Sally told Helga Osk Hlynsdottir when they sat down for the Founders Meeting podcast. A super interesting conversation and inside look of how we work and operate, including having to pivot, competing with bigger AI labs and navigating EU regulation. Worth a listen! https://lnkd.in/evPcx9dv

  • The most interesting thing someone can say to an AI might be nothing at all, such as a short pause, a change in tone, looking away or speaking faster than usual which even though these aren't words, or things we say explicitly they not only affect but sometimes even change the meaning of our words all together. As AI moves from generating text on a screen to tutoring, interviewing, selling, collaborating and operating in the physical world, understanding these signals becomes less of a nice-to-have and more of a requirement, meaning AI will need to understand not just what we communicate, but how we communicate it.

    • Der er ingen alternativ tekst for dette billede
  • "Stubbornness is what helped me succeed." That’s how Paula describes her journey from exchange student in Denmark to leading now leading an AI startup, in this recent profile by The Copenhagen Post. There’s something fitting about that description when you’re trying to build something that doesn’t quite exist yet. You need the curiosity to question what already exists, but also enough stubbornness to keep going when the answer isn’t obvious. A great read on Paula and the story behind Interhuman: https://lnkd.in/eEYv4-wm

  • In AI development text and images came first, and they got the attention, the data, and the breakthroughs, while speech and sound trailed behind as the modality you'd get to eventually. Even now, in a lot of AI systems, audio carries less weight in the actual reasoning than images or text do. But our take is that a system reading the transcript alone knows what you said, but it has no idea what you meant. Consider the sentence: "yeah, that's a great idea." Written down, it's fixed with one meaning on the page. Spoken, it might be genuine enthusiasm, a flat dismissal, or someone too tired to argue, and the only thing that tells those apart is how it was said: the intonation, the pace, the tension in the voice. That's why we treat voice as a core modality rather than an afterthought, and why we brought on people whose whole expertise is the signal hiding inside sound, researchers who've spent years on exactly the part of communication the rest of the field skipped over. Anders Bargum, who came to us straight from a PhD on voice, is one of them.

    • Der er ingen alternativ tekst for dette billede
  • There is a strange paradox in building AI that understands human behaviour: the better we want machines to become at reading people, the more carefully we have to define what we mean by human behaviour in the first place. That is where much of the work happens. Not in the model itself, but in the process of turning something as ambiguous as hesitation, confidence or engagement into data that can actually be learned from. In a new case study with CloudResearch, Alberto Cardaci, our Head of Behavioural Intelligence, goes into some of the thinking behind the data pipeline we have built at Interhuman AI. Across more than 20 multi-wave annotation projects, we have been working with a taxonomy of 12 behavioural signals, continuously calibrating and refining how they are identified. It is a less visible part of building social intelligence, but arguably one of the most important. If the underlying data does not capture behaviour consistently, there is only so much the model can learn from it. CloudResearch takes a closer look at how we approach that problem, and what it takes to build data that can capture something as complex as human behaviour. Read the case study below. https://lnkd.in/ex6asXNf

Tilsvarende sider