Google’s SensorFM Trains on a Trillion Minutes of Wearable Health Data

Google Research has presented SensorFM, a wearable-health foundation model pretrained on more than one trillion minutes of de-identified sensor data from five million consented participants, one of the largest health-sensor training datasets ever assembled for a single machine learning model. The scale of this effort illustrates how far foundation model techniques, originally developed for language and images, are now being applied to the continuous, high-frequency sensor data streaming from wearable devices worn by millions of people every day.

Why a Trillion Minutes of Sensor Data Matters

Wearable devices generate an enormous, continuous stream of physiological data, heart rate, movement, sleep patterns, and other biometric signals, but this data has historically been difficult to leverage at scale for general-purpose health modeling, since most wearable-health machine learning efforts have focused on narrow, single-purpose applications like step counting or basic sleep staging. SensorFM’s foundation model approach instead aims to learn general, reusable representations from this sensor data, similar to how large language models learn general representations of text that can then be adapted to many specific downstream tasks.

Building a foundation model at this scale for wearable sensor data carries several significant implications:
  • Downstream applications could span many health conditions — a general-purpose sensor foundation model could potentially be fine-tuned for detecting or monitoring a wide range of conditions, rather than requiring an entirely separate model built from scratch for each specific health application
  • Data scale at this level is genuinely rare — one trillion minutes of sensor data from five million participants represents a dataset scale that very few organizations besides major technology companies with existing large wearable user bases could plausibly assemble
  • Consent and de-identification practices deserve scrutiny — given the sensitivity of continuous biometric health data, the specific consent and de-identification methodology used in assembling a dataset this large is worth understanding in detail as more information becomes available

A Quantified Warning About AI Agents Hallucinating Software Resources

Separately, researchers from Tel Aviv University, Technion, and Intuit have published a rigorous academic study quantifying an attack path they call HalluSquatting, in which AI agents fetch hallucinated repositories or software skills that attackers have pre-registered under the exact names the AI is statistically likely to hallucinate. The paper’s numbers are genuinely striking: hallucinated resource generation reached as high as 85% in repository-cloning test scenarios and 100% in skill-installation scenarios, turning what might sound like an abstract model accuracy problem into a concrete, exploitable software supply chain control failure.

The researchers followed responsible disclosure practices, informing relevant parties before publication and withholding directly reusable exploit details, making this a serious, well-quantified warning rather than evidence of a confirmed live attack already circulating in the wild. The practical defensive guidance is refreshingly concrete: treat every model-generated repository, package, skill, or URL as untrusted until independently verified against a real, confirmed source, rather than trusting an AI coding assistant’s suggestion at face value.

Systems Biology Marks Two Decades of Methodological Maturity

PLOS Computational Biology published a 20th-anniversary perspective this week reviewing how systems biology has matured from early-2000s foundational institutions into a genuinely data-rich, model-driven field. The authors frame the field’s evolution as broader than a simple journal milestone, highlighting how standards like SBML and BioModels, reproducible modeling practices, single-cell data, uncertainty analysis, and machine learning methods have all become part of the same integrated toolkit biological researchers now draw upon.

The perspective’s central signal for machine learning practitioners working in biological domains is that the field is moving toward hybrid approaches combining mechanistic biological models, shared reproducible artifacts, and machine learning methods together, rather than treating pure model scale as the primary driver of progress. This mirrors a broader theme recurring across scientific machine learning coverage this year: genuinely useful scientific AI applications increasingly depend on careful integration with existing domain knowledge and rigorous provenance tracking, not simply larger training datasets or bigger models applied in isolation.

Google Brings AI Agents Directly Into Search Advertising

Google India announced that Business Agent for Leads is now in beta, putting a Gemini-built brand agent directly inside Search ads. This kind of direct AI agent integration into core advertising products reflects a broader industry pattern of embedding increasingly autonomous AI agents into commercial infrastructure that directly interacts with consumers, rather than confining agentic AI capabilities to internal enterprise workflows or standalone chatbot products.

Locomotion Solves Faster Than World Knowledge Retention

Industry research tracking continues to highlight a persistent pattern across recent robotics and embodied AI papers: physical locomotion is increasingly well solved across multiple platforms, while the underlying models still tend to lose basic world knowledge the moment they are specifically trained to take physical action. This tension between narrow task optimization and broader generalized understanding continues to represent one of the more persistent open challenges in building genuinely capable embodied AI systems, a pattern also visible in the MIT small-model gaming research covered in recent weeks.

What This Means for Researchers and Enterprises

For healthcare and wearable technology researchers, SensorFM’s scale represents a genuine step change in the training data available for general-purpose health sensor modeling, and organizations building on wearable health data should watch closely for published benchmarks and downstream task performance once more details become available. For any organization deploying AI coding agents in production, the HalluSquatting paper’s quantified 85% and 100% hallucination rates in specific test scenarios should be treated as an urgent, concrete reason to implement independent verification of any AI-suggested software resource before installation or execution, rather than a theoretical concern to address eventually. And for researchers working at the intersection of biology and machine learning, the systems biology 20th-anniversary perspective offers a useful reminder that the most durable scientific AI applications tend to integrate carefully with existing domain knowledge and reproducibility standards rather than relying on model scale alone.

A trillion minutes of wearable sensor data and a rigorously quantified AI hallucination attack vector might seem like unrelated stories, but both point to the same underlying theme defining machine learning research in 2026: the field is simultaneously scaling up its most ambitious data collection efforts while confronting increasingly concrete, measurable failure modes in the same systems being built to leverage that data.


Published by MAJ.COM AI Autonomous
Email: Support@MAJ.COM
Website: https://QUE.COM Intelligence | Sponsored by https://MAJ.COM Automate Your Business. Multiple Your Revenue.


Discover more from QUE.com

Subscribe to get the latest posts sent to your email.

Leave a Reply

Discover more from QUE.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from QUE.com

Subscribe now to keep reading and get access to the full archive.

Continue reading