Learning visual to auditory sensory substitution reveals flexibility in image to sound mapping
摘要
Visual-to-auditory sensory substitution devices (SSDs) translate images to sounds. One SSD, The vOICe, translates a pixel’s vertical position into pitch and horizontal position into time. This mapping is primarily based on technical considerations for preserving image content in human-audible sounds without presupposing intuitiveness, although some literature also invokes crossmodal correspondences in perception, such as pitch for elevation. We investigated these presuppositions and the efficacy of learning a traditional algorithm (i.e., pitch indicating elevation and time indicating azimuth) versus a reversed algorithm (i.e., pitch indicating azimuth and time indicating elevation), or an arbitrary single-tone control mapping (i.e., each visual stimulus was represented by a single non-systematic pitch–time pairing without structured spatial correspondences). Sixty sighted adults participated with random assignment to the Traditional, Reversed, or Control groups. They completed learning and evaluation sessions using simplified black-and-white visual stimuli. Both the Traditional and Reversed groups learned mappings within 30 minutes and demonstrated successful recognition of novel stimuli, outperforming the Control group but not differing between them. Structured mappings facilitate SSD learning. Mapping pixel position onto spectral-temporal acoustic axes appears flexible, rather than anchored to cross-modal correspondences. These findings reveal how SSDs may be rendered bespoke across user, stimuli, and functionality levels.