<p>Automated electrographic seizure detection software often rely on the high channel counts of traditional electroencephalography (EEG) recording systems. However, these systems are notoriously cumbersome, limiting both the duration of and access to EEG monitoring. Recent medical-grade wearable devices approach these issues by using small, discreet sensors and reduced channels for ease of use during daily life. The need for reliable seizure detection software that can operate on these reduced-channel recordings will continue to grow as these devices become more widely available. To this end, REMI Vigilenz AI for Event Detection (VED), a novel, reduced-channel automated electrographic seizure detection algorithm, has been developed and commercialized as a clinical decision support tool for one such wearable EEG system (REMI, Epitel, Inc.). As Software as a Medical Device, VED performance was formally assessed against a consensus of expert clinician reviewers of EEG. However, consensus-based approaches, while straightforward to report, can be difficult to interpret, because experts can disagree on what constitutes an electrographic seizure in EEG records. To address this, an inter-rater evaluation paradigm is employed herein, whereby the agreement between the reduced-channel automated detector and clinician experts is directly measured in relation to the degree to which those experts agree among themselves. Additionally, a state-of-the-art, automated electrographic seizure detector designed for high-channel-count EEG (Persyst 15, [P15]) is also assessed to provide further context. To directly simulate the real-world EEG review process, experts and algorithms reviewed entire EEG recordings rather than preselected short-duration snippets. In total, 60 standard-of-care wired EEG records (mean duration: 67 hours) from epilepsy monitoring units and home ambulatory settings, with 19+ channels placed based on the international 10-20 system, were independently annotated for electrographic seizures by groups of three epileptologists (experts) and two algorithms (VED and P15). Experts and P15 reviewed the complete 19+ channel EEG records, while VED operated exclusively on four differential EEG channels extracted from the wired EEG records (equivalent to bilateral frontal and temporoparietal placement as expected by the REMI system). Relative sensitivity, precision, and false positives per day (FPs/day) were computed across expert-expert and algorithm-expert pairs. The experts produced a total of 348 markings across the 4,036 hours of data. Relative inter-rater sensitivity between the experts ranged from 68.4% (95% confidence interval [CI<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(_{95}\)</EquationSource> </InlineEquation>]=[48.8%, 84.5%]) to 88.3%, [77.5%, 94.7%], while relative FPs/day ranged from 0.16, [0.02, 0.47] to 0.79, [0.14, 2.48]. The four-channel VED algorithm averaged a sensitivity of 77.0%, [69.4%-83.2%] with 4.79, [3.83-6.47] FPs/day relative to the experts. P15, provided for context, achieved a relative sensitivity of 65.4%, [56.0%-74.0%] with 1.30, [0.99-1.74] FPs/day. VED non-inferiority to the experts was assessed by setting a -10% sensitivity and +1 FPs/day margin. VED at the official Low confidence level demonstrated non-inferiority in sensitivity (p &lt; 0.01), while not meeting statistical non-inferiority for FPs/day. However, at the Moderate confidence level, VED contributed no more than +0.01 FPs/day relative to experts (p &lt; 0.05) while retaining at least 75% of the experts’ sensitivity (p &lt; 0.01). VED’s sensitivity increased as more expert reviewer agreement was required, even though the median duration of events decreased, demonstrating that inter-rater evaluation methods can provide insights that consensus-based paradigms may misinterpret. The experts had high concordance among themselves and with the site epileptologists when determining which records had electrographic seizure activity. The most discordant records, where VED and experts produced the highest FPs/day, were records where the site physician’s clinical notes stated high rates of polyspike and/or spike-wave activity, epileptiform events that future versions of VED may need to account for. Overall, our findings show that, despite operating on reduced, four-channel EEG data, the VED algorithm achieves inter-rater sensitivity comparable to that of epileptologists and a state-of-the-art full-channel software, albeit at the expense of a higher number of false positives per day. This analysis was only performed on a limited number of participants using conventional wired EEG data rather than <i>in situ</i> wearable EEG data from real-world, everyday life. However, these evaluations lay the groundwork for future inter-rater and consensus-based validation studies that would occur over longer periods of time that will be necessary to demonstrate real-world algorithm performance and clinical utility, including a reduced burden for reviewing extended duration EEG.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing a reduced-channel algorithm for end-to-end seizure detection on multiday EEG using inter-rater agreement with epileptologists

  • Zoë Tosi,
  • Vamshi K. Muvvala,
  • Tyler J. Newton,
  • Avidor B. Kazen,
  • Mark J. Lehmkuhle,
  • Mitchell A. Frankel

摘要

Automated electrographic seizure detection software often rely on the high channel counts of traditional electroencephalography (EEG) recording systems. However, these systems are notoriously cumbersome, limiting both the duration of and access to EEG monitoring. Recent medical-grade wearable devices approach these issues by using small, discreet sensors and reduced channels for ease of use during daily life. The need for reliable seizure detection software that can operate on these reduced-channel recordings will continue to grow as these devices become more widely available. To this end, REMI Vigilenz AI for Event Detection (VED), a novel, reduced-channel automated electrographic seizure detection algorithm, has been developed and commercialized as a clinical decision support tool for one such wearable EEG system (REMI, Epitel, Inc.). As Software as a Medical Device, VED performance was formally assessed against a consensus of expert clinician reviewers of EEG. However, consensus-based approaches, while straightforward to report, can be difficult to interpret, because experts can disagree on what constitutes an electrographic seizure in EEG records. To address this, an inter-rater evaluation paradigm is employed herein, whereby the agreement between the reduced-channel automated detector and clinician experts is directly measured in relation to the degree to which those experts agree among themselves. Additionally, a state-of-the-art, automated electrographic seizure detector designed for high-channel-count EEG (Persyst 15, [P15]) is also assessed to provide further context. To directly simulate the real-world EEG review process, experts and algorithms reviewed entire EEG recordings rather than preselected short-duration snippets. In total, 60 standard-of-care wired EEG records (mean duration: 67 hours) from epilepsy monitoring units and home ambulatory settings, with 19+ channels placed based on the international 10-20 system, were independently annotated for electrographic seizures by groups of three epileptologists (experts) and two algorithms (VED and P15). Experts and P15 reviewed the complete 19+ channel EEG records, while VED operated exclusively on four differential EEG channels extracted from the wired EEG records (equivalent to bilateral frontal and temporoparietal placement as expected by the REMI system). Relative sensitivity, precision, and false positives per day (FPs/day) were computed across expert-expert and algorithm-expert pairs. The experts produced a total of 348 markings across the 4,036 hours of data. Relative inter-rater sensitivity between the experts ranged from 68.4% (95% confidence interval [CI \(_{95}\) ]=[48.8%, 84.5%]) to 88.3%, [77.5%, 94.7%], while relative FPs/day ranged from 0.16, [0.02, 0.47] to 0.79, [0.14, 2.48]. The four-channel VED algorithm averaged a sensitivity of 77.0%, [69.4%-83.2%] with 4.79, [3.83-6.47] FPs/day relative to the experts. P15, provided for context, achieved a relative sensitivity of 65.4%, [56.0%-74.0%] with 1.30, [0.99-1.74] FPs/day. VED non-inferiority to the experts was assessed by setting a -10% sensitivity and +1 FPs/day margin. VED at the official Low confidence level demonstrated non-inferiority in sensitivity (p < 0.01), while not meeting statistical non-inferiority for FPs/day. However, at the Moderate confidence level, VED contributed no more than +0.01 FPs/day relative to experts (p < 0.05) while retaining at least 75% of the experts’ sensitivity (p < 0.01). VED’s sensitivity increased as more expert reviewer agreement was required, even though the median duration of events decreased, demonstrating that inter-rater evaluation methods can provide insights that consensus-based paradigms may misinterpret. The experts had high concordance among themselves and with the site epileptologists when determining which records had electrographic seizure activity. The most discordant records, where VED and experts produced the highest FPs/day, were records where the site physician’s clinical notes stated high rates of polyspike and/or spike-wave activity, epileptiform events that future versions of VED may need to account for. Overall, our findings show that, despite operating on reduced, four-channel EEG data, the VED algorithm achieves inter-rater sensitivity comparable to that of epileptologists and a state-of-the-art full-channel software, albeit at the expense of a higher number of false positives per day. This analysis was only performed on a limited number of participants using conventional wired EEG data rather than in situ wearable EEG data from real-world, everyday life. However, these evaluations lay the groundwork for future inter-rater and consensus-based validation studies that would occur over longer periods of time that will be necessary to demonstrate real-world algorithm performance and clinical utility, including a reduced burden for reviewing extended duration EEG.