Please wait ...
0% Complete
Home
/
11th International Symposium on Telecommunication (IST'2024)
LLM Performance Assessment in Computer Science Graduate Entrance Exams
Authors :
Arya VarastehNezhad
1
Reza Tavasoli
2
Mostafa Masumi
3
Fattaneh Taghiyareh
4
1- University of Tehran
2- University of South Carolina
3- sharif university of technology
4- University of Tehran
Keywords :
Large Language Model،e-Learning،Generative AI،Human-Computer Interaction،Computing Education،LLM Evaluation
Abstract :
Large Language Models (LLMs) are increasingly utilized in educational settings, raising questions about their efficacy in standardized testing contexts. This study evaluates the performance of popular LLMs, including GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Flash, Llama-3.1-70B, Mistral Large 2, DeepSeek-V2, and Gemma-2-27B, in answering data structure and algorithm design questions from the Iranian university entrance exams for master's programs in computer science and information technology. The research analyzes the accuracy of responses, the length of answers, required reading time, and vocabulary levels across the 2022, 2023, and 2024 exams. Questions were categorized into five main areas: Algorithm Analysis and Complexity, Sorting and Searching, Graph Algorithms, Data Structures, and Advanced Topics and NP-completeness. The study compares the LLMs' performance in both Persian and English contexts, providing a comprehensive evaluation of their capabilities and limitations in this domain. Results indicate that GPT-4o achieved the highest average accuracy (75.0%), followed by Claude 3.5 Sonnet (67.2%) and Mistral Large 2 (64.1%). The majority of the models performed better on English questions compared to Persian, with GPT-4o showing the largest performance gap (81.3% vs 68.8%). The findings offer insights into the practical applications of LLMs in educational assessments and contribute to the ongoing discourse on the role of AI in academia, particularly in non-English speaking regions.
Papers List
List of archived papers
Insurance of liability for self-driving cars
Hussein Taleb - Mohammad Hadi Bokaei - Davood Heidari kani - Farbod Rabizadeh Fard - Elham Rafati - Neda Nedaei
Developing 3D CAD Models and Low-Cost Hardware for Learning New Topics in Telecommunication Laboratories
Jamal Kazazi - Mahmoud Kamarei
Trust Analysis Improvement Through Deep Learning for Signed Social Networks
Shayan Karami - Fattaneh Taghiyareh
Exploring Feature Map Correlations for Effective Fake Face Detection
Narges Honarjoo - Fatemeh Taher - Azadeh Mansouri
Synthetic Collision-Prone Trajectory Data Generation Using CTGAN for Connected Autonomous Vehicles
Ali Samanipour - Reza Javidan - Omid Bushehrian
Improving SEO System for Iranian Commercial Website Rankings using Fuzzy Inference System
Yasamin Farjadian - Gholam Ali Montazer - Yeganeh Sattari
Detection of Customer Satisfaction in In-Person Telecommunication Services Using Automated Facial Image Analysis with Artificial Intelligence
Hadiseh ArabGangan - Mehrdad Rohani - Seyed Ali Hosseini - Alireza Mansouri
Deep Reinforcement Learning for AoI-aware Transmission Control in Energy-Harvesting Grant-Free NOMA IoT Networks
Marzieh Sheikhi - Vesal Hakami - Houman Zarrabi
Artificial Neural Network and ARIMA model to estimate ICT sector investment, a comparative approach
Leila Mansourifar - Niloofar Moradhasel - Matin sadat Borghei - Arman Heidari
Proposing a Novel Assessment Framework for Evaluation of Public Cloud Service Providers
Davood Maleki - Ehsan Arianian - Neda Ghorbani
more
Samin Hamayesh - Version 44.9.0