Towards Voice Reconstruction from EEG during Imagined Speech

Young-Eun LeeSeo-Hyun LeeSang-Ho KimSeong-Whan Lee

article2023AAAI62 citations

Presents NeuroTalk, a domain-adaptation framework that synthesizes an individual's own voice directly from non-invasive imagined speech EEG signals and enables phoneme-level reconstruction of unseen words.

Listen

Restoring natural communication for paralyzed or locked-in patients remains a major challenge in assistive technology. While invasive brain-computer interfaces have successfully synthesized audible speech from spoken words, non-invasive methods such as electroencephalography (EEG) face severe obstacles: low signal-to-noise ratios, movement artifacts, and the absence of reference voice recordings during unspoken thoughts.

The article develops and evaluates NeuroTalk, a deep learning framework designed to reconstruct a user's own voice directly from non-invasive EEG signals recorded during imagined speech.

To address the absence of reference audio during silent imagery, the authors introduced a domain adaptation approach. The framework extracts spatial and spectral neural features and projects spoken EEG into the feature space of imagined EEG. The model was pre-trained on paired spoken EEG and vocal audio, then fine-tuned on imagined EEG using generative adversarial network losses, reconstruction losses, and phonetic decoding guidance. The experimental evaluation utilized 64-channel EEG recordings from six participants across 1,300 trials per paradigm over a 12-phrase vocabulary and a silent phase, alongside a leave-one-out test on unseen words.

The article reports several critical findings. First, NeuroTalk successfully synthesized recognizable audio from imagined speech EEG, achieving a subjective mean opinion score of 2.78 out of 5 and a character error rate of 68.3%, compared to 3.34 and 40.2% for spoken EEG. Second, the model accurately detected silent intervals and speech onset, decoding silent phases correctly across nearly all test instances. Third, the framework demonstrated zero-shot capability on unseen words by reconstructing phoneme sequences from pre-trained building blocks, achieving a mean opinion score of 2.57 despite an 83.1% character error rate. Finally, ablation analyses confirmed that reconstruction loss, connectionist temporal classification loss, and domain adaptation are essential components for model accuracy.

These findings indicate that non-invasive, direct brain-to-voice synthesis is technically viable without requiring brain surgery. By bridging spoken and imagined neural patterns, the approach significantly lowers medical risk and equipment burdens for assistive communication interfaces, offering a path toward intuitive communication for individuals who cannot physically speak.

Future development should focus on expanding the vocabulary size beyond isolated words to continuous, sentence-level speech and conducting clinical pilot trials with paralyzed patients. Organizations investing in neural interfaces should continue refining non-invasive domain adaptation pipelines rather than relying exclusively on invasive surgical implants.

Readers should interpret these results with measured caution. The study was conducted in a controlled environment with a small sample of six healthy participants and a constrained vocabulary. While current synthesis performance exhibits moderate error rates, the methodology provides a robust foundation for practical, non-invasive silent communication.

arXiv: 2301.07173
Cover for Towards Voice Reconstruction from EEG during Imagined Speech

Abstract

Translating imagined speech from human brain activity into voice is a challenging and absorbing research issue that can provide new means of human communication via brain signals. Efforts to reconstruct speech from brain activity have shown their potential using invasive measures of spoken speech data, but have faced challenges in reconstructing imagined speech. In this paper, we propose NeuroTalk, which converts non-invasive brain signals of imagined speech into the user’s own voice. Our model was trained with spoken speech EEG which was generalized to adapt to the domain of imagined speech, thus allowing natural correspondence between the imagined speech and the voice as a ground truth. In our framework, an automatic speech recognition decoder contributed to decomposing the phonemes of the generated speech, demonstrating the potential of voice reconstruction from unseen words. Our results imply the potential of speech synthesis from human EEG signals, not only from spoken speech but also from the brain signals of imagined speech.

Table of Contents

  • Introduction
  • Main Contribution
  • Background
  • Speech-Related Paradigms
  • Invasive Approach
  • Non-invasive Approach
  • Method
  • Architectures
  • Training Loss Term
  • Domain Adaptation
  • Dataset
  • Pre-processing
  • Dataset Composition and Training Procedure
  • Experimental Setup
  • Model Implementation Details
  • Evaluation Metrics
  • Results and Discussion Voice Reconstruction from EEG
  • Voice Reconstruction of Unseen Words
  • Ablation Study
  • Domain Adaptation
  • Leave-One-Out Scenario
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — NeuroTalk voice-reconstruction and domain-adaptation framework

    model/method

    NeuroTalk reconstructs a participant’s voice from non-invasive EEG recorded during imagined speech. The pipeline converts EEG into a feature embedding, uses a generator GG to produce a mel-spectrogram, applies a pretrained HiFi-GAN vocoder to synthesize waveform audio, and passes the audio through a pretrained HuBERT-based automatic speech recognition model to obtain text. A discriminator DD distinguishes real from generated mel-spectrograms, while the ASR output supplies a connectionist temporal classification objective that encourages character- or phoneme-level content.

    Because imagined speech has no simultaneously produced reference voice, each imagined-speech trial is paired with the voice recording from the corresponding spoken-speech sequence. Dynamic time warping aligns the generated mel-spectrogram with the mel-spectrogram of that spoken voice. Domain adaptation then transfers information from spoken EEG to imagined EEG in two stages: the imagined-speech CSP subspace is shared with spoken EEG, and a generator–discriminator model trained on spoken EEG and its audio is fine-tuned on imagined EEG using the paired spoken voice as the target. This design aims to preserve the participant’s vocal identity while learning from imagined speech without an imagined-speech audio ground truth.

  2. Knowl 2 — Multi-receptive-field recurrent generator and discriminator

    model/method

    The NeuroTalk generator maps an EEG embedding sequence to a mel-spectrogram using a one-dimensional pre-convolution, a bidirectional gated recurrent unit, repeated upsampling blocks, and a one-dimensional post-convolution followed by a hyperbolic-tangent activation. Each upsampling block uses a transposed one-dimensional convolution and a multi-receptive-field fusion module. An MRF module sums residual-block outputs with different kernel sizes and dilation rates, allowing the generator to represent speech-related temporal patterns at multiple receptive-field scales while residual connections reduce optimization difficulty.

    The discriminator processes mel-spectrograms in the reverse direction using one-dimensional convolutions, MRF modules, downsampling, a recurrent unit, a fully connected layer, and a sigmoid real/fake output. In the experiment, the generator used three residual blocks with kernel sizes 33, 77, and 1111, dilation rates 11, 33, and 55, and upsampling rates 33, 22, and 22. Its initial channel count was 1,0241{,}024, and the directional GRU dimension was half that count. The discriminator used the same residual-block configuration, downsampled by 33 at each of three stages, ended with 6464 channels, and used a directional GRU dimension equal to half of its final channel count. The MRF modules were repeated three times in both networks.

  3. Knowl 3 — Shared-CSP EEG embedding for spoken and imagined speech

    model/method

    Each EEG trial was transformed into an embedding designed to retain spatial and temporal oscillatory information. Common spatial pattern (CSP) filters extracted spatial components, and log-variance features summarized the temporal oscillations. The CSP filters were learned from the imagined-speech training data and then applied unchanged to both imagined-speech EEG and spoken-speech EEG. Sharing filters in this direction maps spoken EEG into the imagined-speech subspace, reducing dependence on mouth-movement and vibration artifacts present during spoken speech.

    A trial contained 5,0005{,}000 time points from 6464 EEG channels. The embedding used eight CSP features and sixteen non-overlapping temporal segments. With thirteen CSP classes and eight features per class, the resulting embedding had shape 104104 features by 1616 time segments. The same shared embedding procedure was used before spoken-speech model training and imagined-speech fine-tuning.

  4. Knowl 4 — Joint reconstruction, adversarial, and CTC training objectives

    equation

    NeuroTalk trains the generator with mel-spectrogram reconstruction, adversarial, and CTC losses, while the discriminator is trained with an adversarial loss. Let ss be an EEG embedding, xx be the target mel-spectrogram paired with ss, G(s)G(s) be the generated mel-spectrogram, and D(z)D(z) be the discriminator’s estimated probability that mel-spectrogram zz is real. The generator and discriminator objectives are

    L(G)=λg1Lrec(G)+λg2Ladv(G;D)+λg3Lctc(G)L(G)=\lambda_{g1}L_{\mathrm{rec}}(G)+\lambda_{g2}L_{\mathrm{adv}}(G;D)+\lambda_{g3}L_{\mathrm{ctc}}(G) L(D)=λdLadv(D;G),L(D)=\lambda_dL_{\mathrm{adv}}(D;G),

    where λg1\lambda_{g1}, λg2\lambda_{g2}, λg3\lambda_{g3}, and λd\lambda_d are loss weights. The reconstruction loss is the expected elementwise squared error over EEG embeddings:

    Lrec(G)=Es[(G(s)−x)2].L_{\mathrm{rec}}(G)=\mathbb{E}_{s}\left[(G(s)-x)^2\right].

    The generator and discriminator adversarial terms are

    Ladv(G;D)=Es[log⁡(1−D(G(s)))]L_{\mathrm{adv}}(G;D)=\mathbb{E}_{s}\left[\log\left(1-D(G(s))\right)\right] Ladv(D;G)=E(x,s)[log⁡(1−D(x))+log⁡(D(G(s)))].L_{\mathrm{adv}}(D;G)=\mathbb{E}_{(x,s)}\left[\log\left(1-D(x)\right)+\log\left(D(G(s))\right)\right].

    The CTC loss is computed from the ASR sequence produced from generated audio and the target text, allowing training without frame-level alignment and encouraging the generator to preserve character- and phoneme-level information useful for unseen words.

  5. Knowl 5 — Six-participant EEG and audio experimental protocol

    experimental setup

    Six participants each produced twelve prompted words or phrases—ambulance, clock, hello, help me, light, pain, stop, thank you, toilet, TV, water, and yes—plus a silent condition in both spoken-speech and imagined-speech paradigms. Each class had 100 trials per paradigm, giving each participant 1,300 spoken-speech trials and 1,300 imagined-speech trials. Spoken EEG and voice were recorded in synchronized trials; the corresponding spoken voice was used as the imagined-speech target because the imagined trials themselves had no voice recording.

    EEG was sampled at 2,5002{,}500 Hz through 64 active Ag/AgCl electrodes, and voice was initially sampled at 8,0008{,}000 Hz. EEG trials were extracted over two seconds, band-pass filtered from 3030 to 120120 Hz with a fifth-order Butterworth filter, notch-filtered at 6060 Hz and its 120120 Hz harmonic, and baseline-corrected using the preceding 500 ms. Blind source separation removed EOG and EMG components from spoken EEG. Voice was resampled to 22,05022{,}050 Hz and denoised.

    The word stop was excluded from training but retained in validation and testing as an unseen-word condition; its phonemes were covered by the other eleven word or phrase classes. Random five-fold train/validation/test partitions were used. Mel-spectrograms used an FFT size and window size of 1,0241{,}024, hop size 256256, and 80 mel bands at 22,05022{,}050 Hz. Initial spoken-speech training used learning rate 10−410^{-4}, imagined-speech fine-tuning used 10−510^{-5}, the maximum training duration was 500 epochs, and the batch size was 10. AdamW used β1=0.8\beta_1=0.8, β2=0.99\beta_2=0.99, weight decay 0.010.01, and an epoch-wise scheduling factor of 0.9990.999.

  6. Knowl 6 — Held-out reconstruction quality for trained word classes

    data/table

    On held-out trials whose word or phrase classes were available during training, NeuroTalk reconstructed spoken EEG more accurately and naturally than imagined EEG. RMSE measures mel-spectrogram error, CER is the character error rate after ASR, and MOS is the 1–5 subjective mean opinion score. The GT-transformed condition is the original ground-truth voice passed through the mel-spectrogram transform and vocoder, providing an upper reference for the synthesis pipeline.

    Model RMSE CER (%) MOS
    GT – 18.4±11.518.4 \pm 11.5 3.67±1.03.67 \pm 1.0
    GT (trans.) – 23.4±10.923.4 \pm 10.9 3.68±0.93.68 \pm 0.9
    SpEEG 0.17±0.020.17 \pm 0.02 40.2±13.540.2 \pm 13.5 3.34±1.03.34 \pm 1.0
    ImEEG 0.18±0.030.18 \pm 0.03 68.3±2.568.3 \pm 2.5 2.78±1.12.78 \pm 1.1

    The spoken-EEG reconstruction achieved MOS close to the transformed ground-truth voice, indicating relatively natural synthesis from spoken EEG. Imagined-EEG reconstruction had higher mel-spectrogram error, higher CER, and lower MOS, but still produced recognizable speech-quality audio in some held-out cases.

  7. Knowl 7 — Phoneme-compositional reconstruction of an unseen word

    empirical result

    NeuroTalk generated the unseen word stop even though stop was excluded from training and its phonemes were learned only from the other eleven word or phrase classes. For spoken EEG, the unseen-word condition achieved RMSE 0.19±0.030.19\pm0.03, CER 78.9±7.4%78.9\pm7.4\%, and MOS 2.87±1.12.87\pm1.1. For imagined EEG, it achieved RMSE 0.19±0.030.19\pm0.03, CER 83.1±14.5%83.1\pm14.5\%, and MOS 2.57±1.22.57\pm1.2.

    The unseen-word audio was inferior to reconstruction of trained classes, but both EEG modalities produced MOS values above 2.52.5, and the CER gap between spoken and imagined EEG was smaller for the unseen word than for trained classes. These results support the possibility that the CTC-guided generator learned reusable character- or phoneme-level information rather than only selecting among fixed training classes. The authors qualify this interpretation because the limited vocabulary also permits a simpler class-confusion explanation.

  8. Knowl 8 — Contribution of recurrent modules, losses, and domain adaptation

    data/table

    An ablation study evaluated imagined-EEG reconstruction by removing one component at a time from the complete NeuroTalk system. The baseline used the recurrent generator and discriminator, GAN loss, reconstruction loss, CTC loss, and the two-stage domain-adaptation procedure. RMSE is mel-spectrogram error, CER is ASR character error rate, and MOS is subjective naturalness.

    Input RMSE CER (%) MOS
    Baseline 0.18±0.030.18 \pm 0.03 68.3±2.568.3 \pm 2.5 2.78±1.12.78 \pm 1.1
    w/o GRU 0.19±0.030.19 \pm 0.03 76.1±3.376.1 \pm 3.3 2.18±1.22.18 \pm 1.2
    w/o GAN loss 0.18±0.020.18 \pm 0.02 76.1±2.376.1 \pm 2.3 2.86±1.22.86 \pm 1.2
    w/o reconstruction loss 0.62±0.120.62 \pm 0.12 80.2±8.080.2 \pm 8.0 2.50±1.32.50 \pm 1.3
    w/o CTC loss 0.39±0.070.39 \pm 0.07 76.9±0.376.9 \pm 0.3 2.52±1.22.52 \pm 1.2
    w/o domain adaptation 0.18±0.030.18 \pm 0.03 72.3±1.772.3 \pm 1.7 2.66±1.22.66 \pm 1.2

    Removing reconstruction loss caused the largest RMSE increase and one of the poorest CER values, making direct mel-spectrogram fidelity the most influential measured component. Removing CTC loss also degraded reconstruction, supporting its role in sequence content. Removing GRUs produced the lowest MOS, indicating that recurrent sequence modeling was particularly important for naturalness. Removing domain adaptation increased CER from 68.3%68.3\% to 72.3%72.3\%, showing that spoken-speech information and the shared EEG subspace contributed to imagined-speech decoding.

  9. Knowl 9 — Silent-speech detection and characteristic imagined-speech failures

    empirical result

    The NeuroTalk system reconstructed silent trials without visible speech activation for spoken EEG and for all but one imagined-EEG case. This indicates that the generator learned both the silent interval and the onset of speech rather than always producing a vocal signal. The result also shows that the shared spoken-to-imagined training strategy can preserve a no-speech intention despite the absence of imagined-speech audio targets.

    Imagined-speech failures were associated with temporal segmentation errors within a phrase. For an example of thank you, a reconstruction with CER 50%50\% produced too little silence between thank and you, while a CER 100%100\% failure generated only a few characters that did not correspond to the target syllables. These cases expose a limitation of the approach: errors in detecting internal pauses can make otherwise partially recovered imagined phrases unintelligible.

  10. Knowl 10 — Preliminary cross-participant transfer and scope limitations

    limitation

    A leave-one-out experiment assessed whether NeuroTalk could extend to a participant whose EEG was excluded from spoken-speech training. For each held-out participant, the model was trained on spoken EEG from the other participants and then fine-tuned using the held-out participant’s imagined EEG. Performance was worse than the ordinary within-participant baseline but better than the version without domain adaptation, providing preliminary evidence for cross-participant transfer.

    This experiment still required imagined-speech data from the held-out participant for fine-tuning and therefore did not demonstrate zero-calibration use by a new person. The study was also based on only six participants, a word-level vocabulary, and relatively poor imagined-speech performance compared with spoken EEG. The authors characterize the results as preliminary and identify larger datasets, more training vocabulary, and sentence-level synthesis as necessary directions for extending the method.

Coverage note — No substantial contributed material was omitted; the framework, EEG representation, architecture, objectives, experimental protocol, quantitative results, unseen-word analysis, ablations, silent decoding, and cross-participant limitations are covered.

References

  1. 1.Akbari, H.; Khalighinejad, B.; Herrero, J. L.; Mehta, A. D.; and Mesgarani, N. 2019. Towards reconstructing intelligible speech from the human auditory cortex. Scientific Reports, 9(1): 1–12.
  2. 2.Angrick, M.; Herff, C.; Mugler, E.; Tate, M. C.; Slutzky, M. W.; Krusienski, D. J.; and Schultz, T. 2019. Speech synthesis from ECoG using densely connected 3D convolutional neural networks. Journal of Neural Engineering, 16(3): 036019.
  3. 3.Angrick, M.; Ottenhoff, M.; Diener, L.; Ivucic, D.; Ivucic, G.; Goulis, S.; Colon, A. J.; Wagner, L.; Krusienski, D. J.; Kubben, P. L.; et al. 2022. Towards closed-loop speech synthesis from stereotactic EEG: A unit delection approach. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1296–1300. IEEE.
  4. 4.Angrick, M.; Ottenhoff, M.; Goulis, S.; Colon, A. J.; Wagner, L.; Krusienski, D. J.; Kubben, P. L.; Schultz, T.; and Herff, C. 2021a. Speech synthesis from stereotactic EEG using an electrode shaft dependent multi-input convolutional neural network approach. In 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 6045–6048. IEEE.
  5. 5.Angrick, M.; Ottenhoff, M. C.; Diener, L.; Ivucic, D.; Ivucic, G.; Goulis, S.; Saal, J.; Colon, A. J.; Wagner, L.; Krusienski, D. J.; et al. 2021b. Real-time synthesis of imagined speech processes from minimally invasive recordings of neural activity. Communications Biology, 4(1): 1–10.
  6. 6.Anumanchipalli, G. K.; Chartier, J.; and Chang, E. F. 2019. Speech synthesis from neural decoding of spoken sentences. Nature, 568(7753): 493–498.
  7. 7.Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33: 12449–12460.
  8. 8.Chaudhary, U.; Birbaumer, N.; and Ramos-Murguialday, A. 2016. Brain–computer interfaces for communication and rehabilitation. Nature Reviews Neurology, 12(9): 513–525.
  9. 9.Cho, K.; Van Merrienboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; and Bengio, Y. 2014. Learning phrase representations using RNN encoder–decoder for statistical machine translation. arXiv preprint arXiv:1406.1078.
  10. 10.Cooney, C.; Folli, R.; and Coyle, D. 2018. Neurolinguistics research advancing development of a direct-speech brain–computer interface. IScience, 8: 103–125.
  11. 11.Delorme, A.; and Makeig, S. 2004. EEGLAB: an open source toolbox for analysis of single-trial EEG dynamics including independent component analysis. Journal of Neuroscience Methods, 134(1): 9–21.
  12. 12.Devlaminck, D.; Wyns, B.; Grosse-Wentrup, M.; Otte, G.; and Santens, P. 2011. Multisubject learning for common spatial patterns in motor-imagery BCI. Computational Intelligence and Neuroscience, 2011.
  13. 13.Gaddy, D.; and Klein, D. 2020. Digital voicing of silent speech. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 5521–5530.
  14. 14.Gaddy, D.; and Klein, D. 2021. An improved model for voicing silent speech. arXiv preprint arXiv:2106.01933.
  15. 15.Goldstein, A.; Zada, Z.; Buchnik, E.; Schain, M.; Price, A.; Aubrey, B.; Nastase, S. A.; Feder, A.; Emanuel, D.; Cohen, A.; et al. 2022. Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3): 369–380.
  16. 16.Gomez-Herrero, G.; De Clercq, W.; Anwar, H.; Kara, O.; Egiazarian, K.; Van Huffel, S.; and Van Paesschen, W. 2006. Automatic removal of ocular artifacts in the EEG without an EOG reference channel. In Proceedings of the 7th Nordic Signal Processing Symposium, 130–133. IEEE.
  17. 17.Gonzalez, J. A.; Cheah, L. A.; Gomez, A. M.; Green, P. D.; Gilbert, J. M.; Ell, S. R.; Moore, R. K.; and Holdsworth, E. 2017. Direct speech reconstruction from articulatory sensor data by machine learning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25(12): 2362–2374.
  18. 18.Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014. Generative adversarial nets. Advances in Neural Information Processing Systems, 27.
  19. 19.Graimann, B.; Allison, B.; and Pfurtscheller, G. 2009. Brain–computer interfaces: A gentle introduction. In Brain–Computer Interfaces, 1–27. Springer.
  20. 20.Graves, A.; Fernandez, S.; Gomez, F.; and Schmidhuber, J. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning, 369–376.
  21. 21.Herff, C.; Diener, L.; Angrick, M.; Mugler, E.; Tate, M. C.; Goldrick, M. A.; Krusienski, D. J.; Slutzky, M. W.; and Schultz, T. 2019. Generating natural, intelligible speech from brain activity in motor, premotor, and inferior frontal cortices. Frontiers in Neuroscience, 13: 1267.
  22. 22.Herff, C.; Krusienski, D. J.; and Kubben, P. 2020. The potential of stereotactic-EEG for brain-computer interfaces: current progress and future directions. Frontiers in Neuroscience, 14: 123.
  23. 23.Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29: 3451–3460.
  24. 24.Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1125–1134.
  25. 25.Kong, J.; Kim, J.; and Bae, J. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33: 17022–17033.
  26. 26.Krepki, R.; Blankertz, B.; Curio, G.; and Muller, K.-R. 2007. The Berlin Brain-Computer Interface (BBCI)–towards a new communication channel for online control in gaming applications. Multimedia Tools and Applications, 33(1): 73–90.
  27. 27.Krishna, G.; Tran, C.; Han, Y.; Carnahan, M.; and Tewfik, A. H. 2020. Speech synthesis using EEG. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1235–1238. IEEE.
  28. 28.Lachaux, J.-P.; Axmacher, N.; Mormann, F.; Halgren, E.; and Crone, N. E. 2012. High-frequency neural activity and human cognition: past, present and possible future of intracranial EEG research. Progress in Neurobiology, 98(3): 279–301.
  29. 29.Lee, M.-H.; Kwon, O.-Y.; Kim, Y.-J.; Kim, H.-K.; Lee, Y.-E.; Williamson, J.; Fazli, S.; and Lee, S.-W. 2019a. EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy. GigaScience, 8(5): giz002.
  30. 30.Lee, S.-H.; Lee, M.; Jeong, J.-H.; and Lee, S.-W. 2019b. Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery. In IEEE International Conference on Systems, Man and Cybernetics (SMC), 4409–4414. IEEE.
  31. 31.Lee, S.-H.; Lee, M.; and Lee, S.-W. 2019. EEG representations of spatial and temporal features in imagined speech and overt speech. In Asian Conference on Pattern Recognition, 387–400.
  32. 32.Lee, S.-H.; Lee, M.; and Lee, S.-W. 2020. Neural decoding of imagined speech and visual imagery as intuitive paradigms for BCI communication. IEEE Transactions on Neural Systems and Rehabilitation Engineering.
  33. 33.Lee, S.-H.; Lee, Y.-E.; and Lee, S.-W. 2022. Toward imagined speech based smart communication system: potential applications on metaverse conditions. In 2022 10th International Winter Conference on Brain-Computer Interface (BCI), 1–4. IEEE.
  34. 34.Loshchilov, I.; and Hutter, F. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  35. 35.Makin, J. G.; Moses, D. A.; and Chang, E. F. 2020. Machine translation of cortical activity to text with an encoder–decoder framework. Nature Neuroscience, 23(4): 575–582.
  36. 36.Meng, K.; Lee, S.-H.; Goodarzy, F.; Vogrin, S.; Cook, M. J.; Lee, S.-W.; and Grayden, D. B. 2022. Evidence of Onset and Sustained Neural Responses to Isolated Phonemes from Intracranial Recordings in a Voice-based Cursor Control Task. In Proceedings of the Annual Conference of the International Speech Communication Association (Interspeech), 4063–4067.
  37. 37.Nguyen, C. H.; Karavas, G. K.; and Artemiadis, P. 2017. Inferring imagined speech using EEG signals: a new approach using Riemannian manifold features. Journal of Neural Engineering, 15(1): 016002.
  38. 38.Proix, T.; Delgado Saa, J.; Christen, A.; Martin, S.; Pasley, B. N.; Knight, R. T.; Tian, X.; Poeppel, D.; Doyle, W. K.; Devinsky, O.; et al. 2022. Imagined speech can be decoded from low-and cross-frequency intracranial EEG features. Nature Communications, 13(1): 1–14.
  39. 39.Saha, P.; Abdul-Mageed, M.; and Fels, S. 2019. SPEAK YOUR MIND! Towards imagined speech recognition with hierarchical deep learning. Proc. Interspeech 2019, 141–145.
  40. 40.Sainburg, T.; Thielk, M.; and Gentner, T. Q. 2020. Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires. PLoS Computational Biology, 16(10): e1008228.
  41. 41.Schultz, T.; Wand, M.; Hueber, T.; Krusienski, D. J.; Herff, C.; and Brumberg, J. S. 2017. Biosignal-based spoken communication: A survey. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25(12): 2257–2271.
  42. 42.Si, X.; Li, S.; Xiang, S.; Yu, J.; and Ming, D. 2021. Imagined speech increases the hemodynamic response and functional connectivity of the dorsal motor cortex. Journal of Neural Engineering, 18(5): 056048.
  43. 43.Wang, L.; Zhang, X.; Zhong, X.; and Zhang, Y. 2013. Analysis and classification of speech imagery EEG for BCI. Biomedical Signal Processing and Control, 8(6): 901–908.
  44. 44.Wang, Z.; and Ji, H. 2021. Open vocabulary Electroencephalography-to-text decoding and zero-shot sentiment classification. arXiv preprint arXiv:2112.02690.
  45. 45.Watanabe, H.; Tanaka, H.; Sakti, S.; and Nakamura, S. 2020. Synchronization between overt speech envelope and EEG oscillations during imagined speech. Neuroscience Research, 153: 48–55.

Citation

MLA
Lee, Y.-E., et al. “Towards Voice Reconstruction from EEG During Imagined Speech”. arXiv, 2023, http://arxiv.org/abs/2301.07173v1.
APA
Lee, Y.-E., Lee, S.-H., Kim, S.-H., & Lee, S.-W. (2023). Towards Voice Reconstruction from EEG during Imagined Speech. arXiv. http://arxiv.org/abs/2301.07173v1
Chicago
Lee, Y.-E., S.-H. Lee, S.-H. Kim, and S.-W. Lee. 2023. “Towards Voice Reconstruction from EEG During Imagined Speech”. arXiv. http://arxiv.org/abs/2301.07173v1.
Harvard
Lee, Y.-E. et al. (2023) “Towards Voice Reconstruction from EEG during Imagined Speech”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.07173v1.
Vancouver
1. Lee Y-E, Lee S-H, Kim S-H, Lee S-W (2023) Towards Voice Reconstruction from EEG during Imagined Speech. arXiv

BibTeX

@article{lee2023towards,
  title = {Towards Voice Reconstruction from EEG during Imagined Speech},
  author = {Lee, Young-Eun and Lee, Seo-Hyun and Kim, Sang-Ho and Lee, Seong-Whan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.07173v1},
  eprint = {2301.07173}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF