<p>We consider the problem of approximating the regression function <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(f_\mu :\, \Omega \rightarrow Y\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi>f</mi> <mi>μ</mi> </msub> <mo>:</mo> <mspace width="0.166667em" /> <mi mathvariant="normal">Ω</mi> <mo stretchy="false">→</mo> <mi>Y</mi> </mrow> </math></EquationSource> </InlineEquation> from noisy <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\mu \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>μ</mi> </math></EquationSource> </InlineEquation>-distributed vector-valued data <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\((\omega _m,y_m)\in \Omega \times Y\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo stretchy="false">(</mo> <msub> <mi>ω</mi> <mi>m</mi> </msub> <mo>,</mo> <msub> <mi>y</mi> <mi>m</mi> </msub> <mo stretchy="false">)</mo> <mo>∈</mo> <mi mathvariant="normal">Ω</mi> <mo>×</mo> <mi>Y</mi> </mrow> </math></EquationSource> </InlineEquation> by an online learning algorithm using a reproducing kernel Hilbert space <i>H</i> (RKHS) as prior. In an online algorithm, i.i.d. samples become available one by one via a random process and are successively processed to build approximations to the regression function. Assuming that the regression function essentially belongs to <i>H</i> (soft learning scenario), we provide estimates for the expected squared error in the RKHS norm of the approximations <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(f^{(m)}\in H\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mi>f</mi> <mrow> <mo stretchy="false">(</mo> <mi>m</mi> <mo stretchy="false">)</mo> </mrow> </msup> <mo>∈</mo> <mi>H</mi> </mrow> </math></EquationSource> </InlineEquation> obtained by a standard regularized online approximation algorithm. In particular, we show an order-optimal estimate <Equation ID="Equ34"> <EquationSource Format="TEX">\( \mathbb {E}(\Vert \epsilon ^{(m)}\Vert _H^2)\le C (m+1)^{-s/(2+s)},\qquad m=1,2,\ldots , \)</EquationSource> <EquationSource Format="MATHML"><math display="block"> <mrow> <mrow> <mi mathvariant="double-struck">E</mi> <mo stretchy="false">(</mo> <mo stretchy="false">‖</mo> </mrow> <msup> <mi>ϵ</mi> <mrow> <mo stretchy="false">(</mo> <mi>m</mi> <mo stretchy="false">)</mo> </mrow> </msup> <mrow> <msubsup> <mo stretchy="false">‖</mo> <mi>H</mi> <mn>2</mn> </msubsup> <mo stretchy="false">)</mo> </mrow> <mo>≤</mo> <mi>C</mi> <msup> <mrow> <mo stretchy="false">(</mo> <mi>m</mi> <mo>+</mo> <mn>1</mn> <mo stretchy="false">)</mo> </mrow> <mrow> <mo>-</mo> <mi>s</mi> <mo stretchy="false">/</mo> <mo stretchy="false">(</mo> <mn>2</mn> <mo>+</mo> <mi>s</mi> <mo stretchy="false">)</mo> </mrow> </msup> <mo>,</mo> <mspace width="2em" /> <mi>m</mi> <mo>=</mo> <mn>1</mn> <mo>,</mo> <mn>2</mn> <mo>,</mo> <mo>…</mo> <mo>,</mo> </mrow> </math></EquationSource> </Equation>where <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\epsilon ^{(m)}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>ϵ</mi> <mrow> <mo stretchy="false">(</mo> <mi>m</mi> <mo stretchy="false">)</mo> </mrow> </msup> </math></EquationSource> </InlineEquation> denotes the error term after <i>m</i> processed data, the parameter <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(0&lt;s\le 1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0</mn> <mo>&lt;</mo> <mi>s</mi> <mo>≤</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation> expresses an additional smoothness assumption on the regression function, and the constant <i>C</i> depends on the variance of the input noise, the smoothness of the regression function, and other parameters of the algorithm. The proof, which is inspired by results on Schwarz iterative methods [<CitationRef CitationID="CR12">12</CitationRef>] in the noiseless case, uses only elementary Hilbert space techniques and minimal assumptions on the noise, the feature map that defines <i>H</i> and the associated covariance operator.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Convergence Analysis of Online Algorithms for Vector-Valued Kernel Regression

  • Michael Griebel,
  • Peter Oswald

摘要

We consider the problem of approximating the regression function \(f_\mu :\, \Omega \rightarrow Y\) f μ : Ω Y from noisy \(\mu \) μ -distributed vector-valued data \((\omega _m,y_m)\in \Omega \times Y\) ( ω m , y m ) Ω × Y by an online learning algorithm using a reproducing kernel Hilbert space H (RKHS) as prior. In an online algorithm, i.i.d. samples become available one by one via a random process and are successively processed to build approximations to the regression function. Assuming that the regression function essentially belongs to H (soft learning scenario), we provide estimates for the expected squared error in the RKHS norm of the approximations \(f^{(m)}\in H\) f ( m ) H obtained by a standard regularized online approximation algorithm. In particular, we show an order-optimal estimate \( \mathbb {E}(\Vert \epsilon ^{(m)}\Vert _H^2)\le C (m+1)^{-s/(2+s)},\qquad m=1,2,\ldots , \) E ( ϵ ( m ) H 2 ) C ( m + 1 ) - s / ( 2 + s ) , m = 1 , 2 , , where \(\epsilon ^{(m)}\) ϵ ( m ) denotes the error term after m processed data, the parameter \(0<s\le 1\) 0 < s 1 expresses an additional smoothness assumption on the regression function, and the constant C depends on the variance of the input noise, the smoothness of the regression function, and other parameters of the algorithm. The proof, which is inspired by results on Schwarz iterative methods [12] in the noiseless case, uses only elementary Hilbert space techniques and minimal assumptions on the noise, the feature map that defines H and the associated covariance operator.