Home

Research
    1. Multimodal XR Micro-Interactions While Walking

Multimodal Interaction
       1. Mondrian’s Symphony:The Harmony of Color, Sound, and Touch
       2. Deep Sea Explorer: A multimodal video game Ubicomp/ISWC 2026*

UX DESIGN
  1. Careshare Kitchenette
  2. M-POS
  3. HuaWei Energy-Efficient Metrics Modal *
  4. Spotify-Re:tasted Daylist
  5. E-Together(Grand Prize)
  6. Fluffy’s Revenge: 2D pixel roguelike game


installation
  1. Photoplantmorph
  2. Language Box


About

Info



I am a HCI researcher and interaction designer with a strong focus on interaction design, UX design, HCI, vR/AR, multimodal XR interaction, media tech and ICT4S.

My expertise lies in research-driven prototyping, user study design, empirical evaluation, and the development of precise and refined interactive systems. I am particularly interested in bridging technical implementation and human-centered research, transforming experimental findings into actionable design insights, product improvements, and context-sensitive interaction solutions.


Read more →

1.Multimodal XR Micro-Interactions While Walking:
“Exploring Unimodal Inputs(Hand, Gaze, Speech and Controller) across brief everyday XR Interactions While Walking”



-Keyu Ren, School of EECS, KTH Royal Institute of Technology, Sweden, keyur@kth.se
-Esteli Mariana Garcia, Chair for Human-Computer Interaction, Technical University of Berlin, Germany, e.garcia@tu-berlin.de
-Ceenu George, Chair for Human-Computer Interaction, Technical University of Berlin, Germany, ceenu.george@tu-berlin.de
-Andrii Matviienko, School of EECS, KTH Royal Institute of Technology, Sweden, andriim@kth.se


Abstract:
This study investigates a multimodal XR input system supporting four interaction modalities—Hand Ray + Pinch, Gaze + Dwell, Speech Commands, and Controller input—for brief everyday interactions while walking. To isolate the characteristics of each modality, the 4 inputs are evaluated independently as unimodal experimental conditions across 3 tasks : notification response, visual search, and PIN-entry.

CCS Concepts: • Human-centered computing → Mixed / augmented reality; • Human-centered computing → Interaction techniques; • Human-centered computing → Empirical studies in HCI; • Human-centered computing → Ubiquitous and mobile computing design and evaluation methods.

Keywords:
Multimodal input; Extended Reality (XR); Mobile interaction; gaze interaction; hand tracking; speech interaction; controller input; brief XR interactions; Meta Quest Pro


thesis project 2026.03-2026.09


Role: Developer & researcher




Introduction

XR is increasingly moving from stationary use toward everyday mobile contexts. During walking, users must divide attention between virtual content, physical navigation, bodily stability, and input control. Although modern headsets support gaze, hand, speech, and controller input, these modalities may respond differently to gait-induced movement, limited attention, fatigue, and environmental constraints. Existing studies often examine one modality, one task, or one design intervention. Less is known about how multiple input modalities compare across different forms of brief everyday XR interaction while users continue walking.

Research question
How do these 4 modality inputs(Hand Ray + Pinch, Gaze + Dwell, Speech, and Controller) affect interaction efficacy, locomotion performance, and user preference across brief everyday XR interactions while walking?

  • Sub-RQ 1 (Safety & Stability): How do specific modality-task pairings differentially impact the user's locomotion stability and walking pace?
  • Sub-RQ 2 (UX & Trade-offs): What are the user trade-offs between physical fatigue, cognitive load, and subjective preferences when applying these modalities to different tasks in a mobile context?


Figure 1: Overview of the multimodal XR prototype, including 4  input modalities and three brief everyday interaction tasks performed while walking


related work

Walking introduces motor–cognitive interference because users must divide attention between locomotion and virtual content. Gait-induced head and body movement can reduce the stability of gaze, hand-ray pointing, and target selection, while users may slow down to preserve interaction accuracy. These constraints suggest that mobile XR interfaces should support low-attention interactions that can be completed without substantially disrupting walking.
Prior mobile-HCI research further shows that short mobile everyday interaction often occurs in a fragmented bursts of attention rather than long continuous sessions. Building on this perspective, this project examines 3 forms of brief everyday XR interaction while walking: notification response, visual target acquisition, and short sequential PIN entry.
Input Modalities in XR
  • Gaze: fast and hands-free, but sensitive to dwell delay and gaze instability;
  • Hand Ray: natural spatial pointing, but affected by jitter, tracking loss and arm fatigue;
  • Speech: hands-free and semantically direct, but affected by recognition latency, noise and social context;
  • Controller: stable discrete input, but requires an external handheld device.


  • system design

    The prototype was developed in Unity for the Meta Quest Pro using video passthrough within a visual field spanning approximately ±15° horizontally and vertically under head-free viewing conditions. It supports eye tracking, optical hand tracking, speech commands, and Touch Pro controller input within one consistent XR environment.
    Tasks DEsign
  • Notification Response: Binary response measured by time;
  • Visual Search: Target acquisition measured by time, errors, spatial distances;
  • PIN Entry: Sequential input measured by time, corrections and WPM;


  • Figure 2: Trial procedures and timing definitions for notification response, visual search, and PIN entry

    Table 1: Interaction Characteristics of the Four Input Modalities

    Table2: Task-Specific Interaction Parameters and Modality Mappings

    Figure 3. Schematic illustration of the relative task-interface placements in the XR prototype


    Interaction techniques:
  • Notification Response: Users selected Accept or Cancel using an 800 ms gaze dwell, hand-ray pointing with pinch confirmation, speech by saying “accept” or “cancel” or controller input by toggling between the two options and pressing the trigger;
  • Visual Search: Users selected the prompted application from a 3 × 4 grid using an 800 ms gaze dwell, hand-ray pointing with pinch confirmation, speech by saying the application name, or controller input by navigating the grid and pressing the trigger;
  • PIN Entry:  Users entered a four-digit PIN using a 500 ms gaze dwell, hand-ray pointing with pinch confirmation, speech by saying digits and enter commands, or controller input by navigating the keypad and pressing the trigger.


  • Figure 4. The 12 modality–task conditions implemented in the prototype


    user study

    The study uses a within-subject, modality-blocked design, which is conducted in a 25 m indoor corridor using Meta Quest Pro video passthrough. Each participant completes all 4 modalities, with task and modality orders counterbalanced to reduce learning and ordering effects and walks back and forth along a straight route at a comfortable self-selected pace. Headset yaw is recorded to identify turnaround intervals.
    Evaluation Metrics:
  • Interaction Efficacy: Completion time, Selection errors, 2D target-distance error, PIN entry time, digits per second (DPS);
  • Locomotion Performance: Baseline walking speed, Pre-interaction speed, Interaction speed, Absolute and relative pace drop, Post-interaction speed offset, Turnaround identification;
  • User Experience:  NASA-TLX, Borg CR10 (muscle fatigue, Eye strain), SUS, Preference ranking, Semi-structured interview;

  • The system would synchronize task-level interaction events with 5 Hz headset-position and yaw data. Walking logs contained two-second pre- and post-interaction windows, consecutive 0.5-second interaction segments, horizontal distance, average speed, acceleration, position, yaw change, and turnaround flags.

    resluts

    A within-subjects study with 28 participants revealed clear task-dependent differences among the 4 input modalities.
    • Interaction efficiency: Controller provided the strongest overall objective performance. For notification response, Controller was fastest (M = 1.84 s), followed by Speech (M = 3.61 s), Gaze + Dwell (M = 3.73 s), and Hand Ray + Pinch (M = 7.38 s). For PIN entry, Controller (M = 10.24 s) and Speech (M = 12.26 s) were substantially faster than Gaze + Dwell (M = 17.24 s) and Hand Ray + Pinch (M = 31.22 s).
    • Accuracy: Visual-search error rates were low across all modalities, with a median error rate of 0% in every condition. No significant modality effect was found, indicating that the main differences concerned speed, workload, and locomotion rather than accuracy.
    • Walking interference: Hand Ray + Pinch caused the greatest disruption to walking, particularly during PIN entry. Participants slowed by an average of 52.79%, from 0.882 m/s before interaction to 0.437 m/s during input, and showed incomplete post-interaction recovery. Speech and Controller produced comparatively small walking-speed changes.
    • Workload, fatigue, and usability: Hand Ray + Pinch produced the highest physical demand, effort, frustration, arm fatigue, and finger fatigue. Gaze + Dwell produced the greatest eye strain. Speech received the highest SUS score (83.48), followed by Controller (76.61), Gaze + Dwell (68.48), and Hand Ray + Pinch (50.36).

    Speech was ranked first overall by 42.9% of participants and Controller by 35.7%. Interviews supported these patterns: Speech was considered effortless and hands-free, Controller stable and accurate, Gaze convenient but visually demanding, and Hand Ray + Pinch natural but difficult to stabilise while walking.




    discussion

    The results demonstrate that there is no universally superior XR input modality while walking. Instead, modality suitability depends on task length, target characteristics, interaction sequence, and surrounding context.
    • Speech offered the strongest overall hands-free balance in the tested quiet indoor environment. It combined high usability and preference with low physical demand and limited walking interference. However, its suitability may decrease in noisy, public, privacy-sensitive, or socially uncomfortable situations.
    • Controller input provided the most reliable objective performance. Its tactile controls supported accurate and stable interaction during walking, but the need to hold an external device limits its usefulness when the user is carrying other objects or requires both hands to remain available.
    • Gaze + Dwell was effective for brief, single-selection interactions, such as responding to a notification. Its fixed dwell duration and sustained fixation became increasingly costly during repeated PIN selections, producing slower input and greater eye strain. Gaze may therefore be more suitable for target specification than repeated confirmation.
    • Hand Ray + Pinch provided device-free and spatially direct interaction, but walking-induced instability, pinch timing, and unsupported arm movement generated substantial performance, fatigue, and locomotion costs. It appears better suited to large targets and short action sequences than to precise or repetitive input.

    Walking-speed reductions should be interpreted as locomotion interference rather than direct evidence of accident risk. Participants may also have slowed deliberately to maintain balance or interaction accuracy. Overall, the findings support task-aware and adaptive XR interfaces that select or combine modalities according to interaction demands and environmental conditions.

    Selected references

    [13] Y. Shin, A. Esteves, and I. Oakley, “Using augmented reality on the go: Understanding the effects of mobility on user performance and subjective workload across eye, head, and hand ray pointing,” International Journal of Human-Computer Studies, vol. 204, Art. no. 103597, 2025. https://doi.org/10.1016/j.ijhcs.2025.103597
    [18] T. Li, E. Velloso, A. Withana, and Z. Sarsenbayeva, “Estimating the effects of encumbrance and walking on mixed reality interaction,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, Art. no. 1153, pp. 1–24, https://dl.acm.org/doi/10.1145/3706598.3713492
    [27] M. N. Lystbæk, K. Pfeuffer, J. E. S. Grønbæk, and H. Gellersen, “Exploring gaze for assisting freehand selection-based text entry in AR,” Proceedings of the ACM on Human-Computer Interaction, vol. 6, no. ETRA, 2022. https://doi.org/10.1145/3530882
    [39] J. D. Hincapié-Ramos, X. Guo, P. Moghadasian, and P. Irani, “Consumed endurance: A metric to quantify arm fatigue of mid-air interactions,” in Proceedings of CHI ’14, 2014. https://doi.org/10.1145/2556288.2557130
    [60] F. Alallah et al., “Performer vs. observer: Whose comfort level should we consider when examining the social acceptability of input modalities for head-worn display?” in Proceedings of VRST ’18, 2018. https://doi.org/10.1145/3281505.3281541
    [64] J. K. Adhikary and K. Vertanen, “Text entry in virtual environments using speech and a midair keyboard,” IEEE Transactions on Visualization and Computer Graphics, vol. 27, no. 5, pp. 2648–2658, 2021. https://doi.org/10.1109/TVCG.2021.3067776
    [68] A. Matviienko, J.-B. Durand-Pierre, J. Cvancar, and M. Mühlhäuser, “Text me if you can: Investigating text input methods for cyclists,” in CHI ’23 Extended Abstracts, 2023. https://doi.org/10.1145/3544549.3585734
    [72] F. Kern, F. Niebling, and M. E. Latoschik, “Text input for non-stationary XR workspaces: Investigating tap and word-gesture keyboards,” IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 5, pp. 2658–2669, 2023. https://doi.org/10.1109/TVCG.2023.3247098
    [76] L. Plabst, F. Niebling, S. Oberdörfer, and F. R. Ortega, “Order up! Multimodal interaction techniques for notifications in augmented reality,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 5, pp. 2258–2267, 2025. https://doi.org/10.1109/TVCG.2025.3549186
    [77] H. Lee and W. Woo, “Exploring the effects of augmented reality notification type and placement in AR HMD while walking,” in Proceedings of IEEE VR 2023, pp. 519–529, 2023. https://doi.org/10.1109/VR55154.2023.00067

    More detailes in the forthcoming thesis and a potential publication