Dataconomy
  • News
    • Artificial Intelligence
    • Cybersecurity
    • DeFi & Blockchain
    • Finance
    • Gaming
    • Startups
    • Tech
  • Industry
  • Research
  • Resources
    • Articles
    • Guides
    • Case Studies
    • Whitepapers
    • AI Models Leaderboard
  • AI toolsNEW
  • Newsletter
  • + More
    • Glossary
    • Conversations
    • Events
    • About
      • Who we are
      • Contact
      • Imprint
      • Legal & Privacy
      • Partner With Us
Subscribe
No Result
View All Result
  • AI
  • Tech
  • Cybersecurity
  • Finance
  • DeFi & Blockchain
  • Startups
  • Gaming
Dataconomy
  • News
    • Artificial Intelligence
    • Cybersecurity
    • DeFi & Blockchain
    • Finance
    • Gaming
    • Startups
    • Tech
  • Industry
  • Research
  • Resources
    • Articles
    • Guides
    • Case Studies
    • Whitepapers
    • AI Models Leaderboard
  • AI toolsNEW
  • Newsletter
  • + More
    • Glossary
    • Conversations
    • Events
    • About
      • Who we are
      • Contact
      • Imprint
      • Legal & Privacy
      • Partner With Us
Subscribe
No Result
View All Result
Dataconomy
No Result
View All Result

New Apple paper reveals how AI can track your daily chores

The team utilized the Ego4D dataset to analyze twelve distinct activities like cooking and cycling.

byKerem Gülen
November 23, 2025
in Research
Home Research
Share on FacebookShare on TwitterShare on LinkedInShare on WhatsAppShare on e-mail
Google Preferred Source

Apple researchers published a study detailing how large language models (LLMs) can interpret audio and motion data to identify user activities, focusing on late multimodal sensor fusion for activity recognition.

The paper, titled “Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition,” by Ilker Demirel, Karan Ketankumar Thakkar, Benjamin Elizalde, Miquel Espi Marques, Shirley Ren, and Jaya Narain, was accepted at the Learning from Time Series for Health workshop at NeurIPS 2025. This research explores integrating LLM analysis with traditional sensor data to enhance activity classification.

The researchers state, “Sensor data streams provide valuable information around activities and context for downstream applications, though integrating complementary information can be challenging. We show that large language models (LLMs) can be used for late fusion for activity classification from audio and motion time series data.” They curated a subset of data for diverse activity recognition from the Ego4D dataset, encompassing household activities and sports.

Stay Ahead of the Curve!

Don't miss out on the latest insights, trends, and analysis in the world of data, technology, and startups. Subscribe to our newsletter and get exclusive content delivered straight to your inbox.

Evaluated LLMs achieved 12-class zero- and one-shot classification F1-scores significantly above chance, without task-specific training. Zero-shot classification through LLM-based fusion from modality-specific models enables multimodal temporal applications with limited aligned training data for a shared embedding space. LLM-based fusion allows model deployment without requiring additional memory and computation for targeted application-specific multimodal models.

The study highlights LLMs’ ability to infer user activities from basic audio and motion signals, showing improved accuracy with a single example. Crucially, the LLM was not directly fed raw audio. Instead, it received short text descriptions generated by audio models and an IMU-based motion model, which tracks movement via accelerometer and gyroscope data.

For the study, researchers utilized Ego4D, a dataset featuring thousands of hours of first-person perspective media. They curated a dataset of daily activities from Ego4D by searching narrative descriptions. The curated dataset includes 20-second samples from twelve high-level activities:

These activities were chosen to cover household and fitness tasks and based on their prevalence in the larger Ego4D dataset. Audio and motion data were processed through smaller models to generate text captions and class predictions. These outputs were then fed into different LLMs, specifically Gemini-2.5-pro and Qwen-32B, to assess activity identification accuracy.

Apple compared model performance in two scenarios: a closed-set test where models chose from the 12 predefined activities, and an open-ended test without provided options. Various combinations of audio captions, audio labels, IMU activity prediction data, and extra context were used for each test.

The researchers noted that the results offer insights into combining multiple models for activity and health data. This approach is particularly beneficial when raw sensor data alone is insufficient to provide a clear picture of user activity. Apple also published supplemental materials, including Ego4D segment IDs, timestamps, prompts, and one-shot examples, to facilitate reproducibility for other researchers.


Featured image credit

Tags: AIAppleego4d

Related Posts

Startup unveils AI model built on oscillators and it could cut energy use by 1,000x

Startup unveils AI model built on oscillators and it could cut energy use by 1,000x

July 21, 2026
Digital transformation of procurement processes: Building a corporate procurement system based on the example of an international industrial holding project

Digital transformation of procurement processes: Building a corporate procurement system based on the example of an international industrial holding project

July 16, 2026
New dark matter theory proposes two particle types

New dark matter theory proposes two particle types

July 14, 2026
Google Dialogflow CX flaw let researchers create rogue agents

Google Dialogflow CX flaw let researchers create rogue agents

July 14, 2026
Penn State researchers build battery-free solar computing chip

Penn State researchers build battery-free solar computing chip

July 14, 2026
Anthropic research introduces GRAM for isolating dangerous AI knowledge

Anthropic research introduces GRAM for isolating dangerous AI knowledge

July 9, 2026

LATEST NEWS

X releases redesigned Android app with faster performance

Google reportedly develops Frozen v2 chip for Gemini AI

Samsung Galaxy Watch Ultra 2 renders leak

NVIDIA unveils hot-water cooled AI servers

Amazon rolls out Adaptive Display for Fire TV

Moonshot pauses Kimi K3 signups amid GPU shortage

BEST AI MODELS LEADERBOARD

See the best AI models, ranked by intelligence, benchmark results, speed and token price. Find the most suitable LLMs, Text-to-Image, Image Editing, Text-to-Speech, Text-to-Video and Image-to-Video  artificial intelligence model for your tasks and business.

LATEST TOOLS

Amanda AI

InterviewBot

VernAI

MyLoans

Essay Grader AI

Cover Letter AI

Animate Old Photos

Resume.io

MonAI

AIEngine Plugin

Dataconomy

COPYRIGHT © DATACONOMY MEDIA GMBH, ALL RIGHTS RESERVED.

  • About
  • Imprint
  • Contact
  • Legal & Privacy

Follow Us

  • News
    • Artificial Intelligence
    • Cybersecurity
    • DeFi & Blockchain
    • Finance
    • Gaming
    • Startups
    • Tech
  • Industry
  • Research
  • Resources
    • Articles
    • Guides
    • Case Studies
    • Whitepapers
    • AI Models Leaderboard
  • AI tools
  • Newsletter
  • + More
    • Glossary
    • Conversations
    • Events
    • About
      • Who we are
      • Contact
      • Imprint
      • Legal & Privacy
      • Partner With Us
No Result
View All Result
Subscribe

This website uses cookies to improve your experience. You can choose to accept or reject them. Visit our Privacy Policy.