Self-supervised pre-training of transformer-based models has demonstrated strong potential for human activity recognition (HAR) using wearable inertial measurement units (IMUs). The existing approaches treat the multi-channel sensor signal as a flat input, disregarding the physical complementarity between accelerometer and gyroscope modalities. In this paper, we propose a correlation-aware tokenization pipeline that explicitly exploits the structural relationship between the two sensor types to construct a compact but semantically rich vocabulary for Large Language Models (LLMs) pre-training. A dual-branch autoencoder processes accelerometer and gyroscope streams independently through specific encoders, then integrates their representations via a cross-modal fusion mechanism that captures temporal correlations. The resulting latent embeddings are subsequently discretized through k-means clustering to produce a fixed vocabulary of motion tokens. We evaluate the proposed tokenizer on four benchmark HAR datasets: RealWorld2016, MHEALTH, WISDM and MotionSense, measuring vocabulary quality through cluster separability metrics and reconstruction error. Experimental results demonstrate that correlation-aware tokenization yields a more compact vocabulary than flat multi-channel baselines while maintaining the information content. These findings suggest that encoding domain-specific physical priors into the tokenization stage constitutes an effective approach for self-supervised learning on IMU data.

Tokenization of IMU Signals for LLMs-Based Human Activity Recognition

Leonardo Bigelli;Emanuele Lattanzi
In corso di stampa

Abstract

Self-supervised pre-training of transformer-based models has demonstrated strong potential for human activity recognition (HAR) using wearable inertial measurement units (IMUs). The existing approaches treat the multi-channel sensor signal as a flat input, disregarding the physical complementarity between accelerometer and gyroscope modalities. In this paper, we propose a correlation-aware tokenization pipeline that explicitly exploits the structural relationship between the two sensor types to construct a compact but semantically rich vocabulary for Large Language Models (LLMs) pre-training. A dual-branch autoencoder processes accelerometer and gyroscope streams independently through specific encoders, then integrates their representations via a cross-modal fusion mechanism that captures temporal correlations. The resulting latent embeddings are subsequently discretized through k-means clustering to produce a fixed vocabulary of motion tokens. We evaluate the proposed tokenizer on four benchmark HAR datasets: RealWorld2016, MHEALTH, WISDM and MotionSense, measuring vocabulary quality through cluster separability metrics and reconstruction error. Experimental results demonstrate that correlation-aware tokenization yields a more compact vocabulary than flat multi-channel baselines while maintaining the information content. These findings suggest that encoding domain-specific physical priors into the tokenization stage constitutes an effective approach for self-supervised learning on IMU data.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11576/2782414
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact