<p>Extracting structured data from web pages has long been a challenging research topic, particularly for domain-specific data on vertical websites. The advent of large language models (LLMs) has introduced new possibilities for this task, enabling high-accuracy information extraction from individual pages. However, due to the substantial costs associated with LLMs, this approach is often economically impractical and difficult to scale in industrial applications. To address these challenges we propose a new framework that leverages a multi-task decomposition strategy, enabling LLMs, despite their limitations in handling structured information, to learn robust XPath expressions from seed pages and apply them to unseen web pages. Experimental results on various structured web data extraction datasets demonstrate that this framework achieves state-of-the-art accuracy for vertical websites with minimal LLM interactions. Furthermore, our approach significantly outperforms the top prompting-based methods, Reflexion and COT, improving F1 scores ranging from 14.84 to 23.86 percentage points across multiple datasets. The proposed framework not only boosts the accuracy and efficiency of structured data extraction from vertical websites, but also holds promise for adaptation across a range of industrial applications focused on extracting information from hierarchical data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic XPath generation agents for vertical websites by LLMs

  • Jing Huang,
  • Jie Song

摘要

Extracting structured data from web pages has long been a challenging research topic, particularly for domain-specific data on vertical websites. The advent of large language models (LLMs) has introduced new possibilities for this task, enabling high-accuracy information extraction from individual pages. However, due to the substantial costs associated with LLMs, this approach is often economically impractical and difficult to scale in industrial applications. To address these challenges we propose a new framework that leverages a multi-task decomposition strategy, enabling LLMs, despite their limitations in handling structured information, to learn robust XPath expressions from seed pages and apply them to unseen web pages. Experimental results on various structured web data extraction datasets demonstrate that this framework achieves state-of-the-art accuracy for vertical websites with minimal LLM interactions. Furthermore, our approach significantly outperforms the top prompting-based methods, Reflexion and COT, improving F1 scores ranging from 14.84 to 23.86 percentage points across multiple datasets. The proposed framework not only boosts the accuracy and efficiency of structured data extraction from vertical websites, but also holds promise for adaptation across a range of industrial applications focused on extracting information from hierarchical data.