Close

Presentation

Evaluating the Impact of System Prompt Length on the Consistency and Accuracy of ChatGPT-Generated Educational Tasks
DescriptionThis study investigates the impact of system prompt length on the consistency and accuracy of ChatGPT-generated educational tasks in an introductory psychology course. A custom OpenAI GPT model was used to teach the Fundamental Attribution Error (FAE), but students reported inconsistencies in assignment delivery. To evaluate the model’s reliability, we developed a rubric assessing (1) adherence to prompt structure, (2) response accuracy, and (3) overall interaction quality. In a classroom pilot (N = 74), feedback informed prompt refinements, and a subsequent study randomly assigned students (N = 259) to either a long (720-word) or short (316-word) system prompt condition. Raters will evaluate AI-student transcripts to determine whether system prompt length influenced instructional fidelity. Preliminary results suggest that longer prompts may produce more accurate and structured responses, while shorter prompts may lead to greater variability. This study aims to clarify whether prompt design significantly affects the reliability of LLM-driven instruction. Findings will inform best practices for using AI in education, emphasizing the need for validated design frameworks to ensure consistent learning outcomes. As AI becomes more prevalent in classrooms, understanding how to optimize prompt construction is critical for achieving scalable, high-quality instructional experiences.