Thalamus AI Secures Millions in Funding for Native Multimodal Long-Term Memory

One-Sentence Summary
Thalamus AI — the only company in China developing native multimodal long-term memory — has launched MemAura, its foundational multimodal memory base, betting on AI’s evolution from general-purpose → personalized → proactive intelligence.
What Is Proactive Intelligence?
Proactive intelligence means AI anticipates user needs and initiates timely, context-aware interactions — after deeply understanding the user over time. To achieve this, robust, cross-modal, persistent memory is non-negotiable.
Funding Round
✅ Completed: Multi-million-dollar seed round
✅ Investors: Shenzhen-based leading fund + strategic industrial capital
✅ Use of Funds: Core R&D acceleration and elite talent acquisition
Founding Team
- Founded: November 2025
- CEO & Founder: Zhang Yuan — dual-degree graduate from Peking University (Electronics & Economics); former COO at an autonomous driving startup; background in venture capital and embodied AI.
- Team Profile: Avg. age ~26; core members from Alibaba DAMO Academy, Tencent, SenseTime, CUHK, HKUST, Peking University, Xi’an Jiaotong University, KIT (Germany), and Fudan University.
Product & Business: MemAura Memory Base
🌟 Key Innovation
Unlike conventional text-only memory systems, MemAura is natively multimodal — designed from day one to ingest, encode, and recall visual, audio, textual, and temporal signals without forcing them into text-first pipelines.
⚙️ Technical Architecture
MemAura implements a biomimetic three-stage memory pipeline:
1. Event Recognition: Filters noise and retains salient semantic & perceptual anchors (like human hippocampal encoding).
2. Multimodal Encoding: Embeds heterogeneous inputs into aligned, time-aware latent representations.
3. Privacy-Aware Partitioning: Separates sensitive vs. public memories with strict access controls — plus hot/cold storage: low-latency retrieval for frequent evidence; cost-efficient archival for long-tail data.
📈 Performance Benchmarks
| Metric | MemAura Result |
|---|---|
| Input token reduction | ↓ 40–49% |
| Memory retrieval latency | < 400 ms |
| End-to-end first-response time | < 1 second |
| Cross-modal recall accuracy | > 80% |
🧩 Deployment Models
- ADK (Agent Development Kit): Pre-packaged memory modules tailored to vertical scenarios (e.g., child companionship vs. adult digital employees).
- APIs: Plug-and-play memory services enabling LLMs and agents to dynamically retrieve and update user state.
- Pricing: Tiered by API call volume + scenario-specific capability licensing.
Industry Context & Competitive Edge
🔍 Market Gap
Current memory solutions suffer from three critical flaws:
– Costly inference: Dumping raw context into oversized windows inflates compute spend.
– Session fragmentation: No continuity across tasks or conversations.
– Modality poverty: Text-only handling ignores vision, speech, sensor streams — blocking true embodied interaction.
Additionally, Transformer-based models inherently struggle with temporal grounding, causing chronological drift and unreliable long-horizon reasoning.
🏆 Benchmark Leadership
In May 2026, Thalamus AI co-launched MEMLENS — the world’s first open multimodal long-memory benchmark — with NVIDIA, HKUST, and CUHK.
🔍 Key findings from MEMLENS (tested on 27 VLMs + 7 Memory Agents):
– Pure-context models degrade sharply beyond 100K tokens; most Memory Agents lose >30% visual fidelity during ingestion.
– Longer contexts increase hallucination — especially when evidence is sparse.
– Post-training memory enhancements often erode model refusal capability, raising safety risks.
✅ Conclusion: Base models (VLMs) and memory layers must specialize — VLMs optimize perception & alignment; memory infrastructures own cross-modal retrieval, state maintenance, and dynamic evidence orchestration.
Core Technical Evolution
| Generation | Architecture | Key Strength | Limitation Addressed |
|---|---|---|---|
| v1 | STKG (Spatio-Temporal Knowledge Graph) | SOTA on LoCoMo & LongMemEval | High engineering overhead; latency-sensitive |
| v2 | MemAura (Hippocampal-Cortical Biomimetic) | Sub-400ms latency, 40%+ token savings | Real-world responsiveness & cost efficiency |
| v3 (in R&D) | E2P (Embedding-to-Prefix) + Transformer micro-tuning | Near-continuous preference learning; aggressive token compression | Adaptive long-term personalization without retraining |
Founder Vision: Five Pillars of Memory Intelligence
🔹 Personalization → Proactivity
“Agents won’t end at task-driven execution — they’ll evolve into state-driven planners. Future agents will self-initiate actions based on accumulated memory.”
🔹 Memory ≠ Storage — Memory = Decision Engine
It decides what to store, how to index it, when to retrieve, and whether to proactively surface insights — turning passive archives into active cognitive infrastructure.
🔹 Long Context ≠ Memory
DeepSeek-V4’s 1M-token window improves KV caching — but true memory requires semantic persistence, cross-session coherence, and multimodal grounding — none of which are solved by scale alone.
🔹 Embodied AI Demands Multimodal Memory First
Home robots, wearables, and ambient agents interact via voice, gesture, gaze, and environment — demanding memory that lives in time and space, not just text buffers.
🔹 Time Is the Anchor
“Time is singular and irreplaceable. All memory must be anchored to it — no abstraction can bypass temporal integrity.”
Early Adoption & Roadmap
- Current Clients: High-volume companion hardware (e.g., caregiving robots), vertical agents (AI customer service, digital staff).
- Next Horizon: Expansion into multimodal enterprise workflows (e.g., medical scribe agents, construction site supervisors) and real-time embodied applications.
- Open Ecosystem: MEMLENS is fully open-sourced on Hugging Face (Top-3 Daily Paper); MemAura SDKs available for early partners.
Article sourced from “Intelligent Emergence”, author Wang Xinyi.