Self-supervised pre-training of transformer-based models has demonstrated strong potential for human activity recognition (HAR) using wearable inertial measurement units (IMUs). The existing approaches treat the multi-channel sensor signal as a flat input, disregarding the physical complementarity between accelerometer and gyroscope modalities. In this paper, we propose a correlation-aware tokenization pipeline that explicitly exploits the structural relationship between the two sensor types to construct a compact but semantically rich vocabulary for Large Language Models (LLMs) pre-training. A dual-branch autoencoder processes accelerometer and gyroscope streams independently through specific encoders, then integrates their representations via a cross-modal fusion mechanism that captures temporal correlations. The resulting latent embeddings are subsequently discretized through k-means clustering to produce a fixed vocabulary of motion tokens. We evaluate the proposed tokenizer on four benchmark HAR datasets: RealWorld2016, MHEALTH, WISDM and MotionSense, measuring vocabulary quality through cluster separability metrics and reconstruction error. Experimental results demonstrate that correlation-aware tokenization yields a more compact vocabulary than flat multi-channel baselines while maintaining the information content. These findings suggest that encoding domain-specific physical priors into the tokenization stage constitutes an effective approach for self-supervised learning on IMU data.
Tokenization of IMU Signals for LLMs-Based Human Activity Recognition
Leonardo Bigelli;Emanuele Lattanzi
In corso di stampa
Abstract
Self-supervised pre-training of transformer-based models has demonstrated strong potential for human activity recognition (HAR) using wearable inertial measurement units (IMUs). The existing approaches treat the multi-channel sensor signal as a flat input, disregarding the physical complementarity between accelerometer and gyroscope modalities. In this paper, we propose a correlation-aware tokenization pipeline that explicitly exploits the structural relationship between the two sensor types to construct a compact but semantically rich vocabulary for Large Language Models (LLMs) pre-training. A dual-branch autoencoder processes accelerometer and gyroscope streams independently through specific encoders, then integrates their representations via a cross-modal fusion mechanism that captures temporal correlations. The resulting latent embeddings are subsequently discretized through k-means clustering to produce a fixed vocabulary of motion tokens. We evaluate the proposed tokenizer on four benchmark HAR datasets: RealWorld2016, MHEALTH, WISDM and MotionSense, measuring vocabulary quality through cluster separability metrics and reconstruction error. Experimental results demonstrate that correlation-aware tokenization yields a more compact vocabulary than flat multi-channel baselines while maintaining the information content. These findings suggest that encoding domain-specific physical priors into the tokenization stage constitutes an effective approach for self-supervised learning on IMU data.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


