Evaluating the Behavior of Small Language Models in Answering Binary Questions
摘要
Small Language Models (SLMs) offer efficient and accessible solutions for natural language processing. This study evaluates their performance in binary question-answering tasks, focusing on prompt sensitivity, multilingual disparities, and token probability analysis. Experiments with Llama-3.2, Mistral-7B, and Phi-3.5-mini reveal that True-First prompts increase TRUE token probability by an average of 0.4, while underrepresented languages like Afrikaans and Polish exhibit notable performance gaps. Token probability analysis uncovers biases toward affirmative responses and cross-lingual challenges, highlighting the critical role of prompt design, diverse training data, and inclusive evaluation. These insights demonstrate the potential of SLMs for cost-effective applications in education, fact-checking, and multilingual NLP, particularly in low-resource settings.