How Does ChatGPT Do on a College Level Astrophysics Exam?

How Does ChatGPT Do on a College Level Astrophysics Exam?

🎙 David Kipping 👥 1.1M 📅 January 7, 2023 ⏱ 28 min 👁 355K 📄 original study 🧭 2026-08-26
Available in: English (current) Français

Keywords

ChatGPTastrophysicsexamAIeducation

Summary

In this video, astrophysicist David Kipping tests ChatGPT’s performance on a final exam from his introductory astronomy course ‘Another Earth’ at Columbia University. The exam consists of multiple-choice and short-answer questions covering topics like planet formation, the Stefan-Boltzmann law, Kepler’s laws, tidal forces, and exoplanet transits. Kipping feeds the questions to ChatGPT and grades its responses using the same rubric as for his human students. ChatGPT scores 73.9%, slightly below the median human score of 75.6%. The video highlights ChatGPT’s strengths in conceptual understanding and its weaknesses in mathematical reasoning and unit handling. Kipping also discusses the implications of AI for education, emphasizing the need to adapt teaching and assessment methods. The video includes a sponsorship segment for Ground News.

120 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable empirical data on ChatGPT’s capabilities in a specialized academic context. The argumentation is solid: Kipping clearly explains each question, the expected reasoning, and why ChatGPT’s answers are correct or incorrect. He also compares ChatGPT’s performance to that of his students, providing a meaningful benchmark. The discussion of potential biases and limitations of the test (e.g., the inability to input graphs) adds nuance. The video’s value lies in its practical demonstration and thoughtful analysis of AI’s role in education.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates scientific rigor by using a real exam, transparent grading, and clear explanations. Kipping references relevant concepts (e.g., Stefan-Boltzmann law, Kepler’s laws) and mentions prior work (e.g., Lanusse’s generative networks). However, he does not cite specific academic sources for these claims, relying on general knowledge. The title accurately reflects the content. The video includes a sponsorship segment for Ground News, which is clearly disclosed. The analysis of comments shows a generally positive reception, with viewers appreciating the educational content and the balanced perspective on AI.

184 words

Title / Content Match

The title accurately reflects the content: a systematic test of ChatGPT on a college-level astrophysics exam.

Quality & Reliability

8/10

The video presents a well-structured, transparent evaluation of ChatGPT's performance on a real exam, with clear grading criteria and comparison to human students. The methodology is sound, though limited to a single exam and subject.

Chapters

Cited Sources

Concurring Sources

  • Ground News — Sponsor, but also a source for media bias information.

Contribution & Novelties

This video offers a novel, hands-on evaluation of ChatGPT’s performance on a real college exam, providing concrete data on its strengths and weaknesses. It contributes to the ongoing discussion about AI’s impact on education by demonstrating both its potential and its current limitations. The comparison with human student performance adds a valuable benchmark.

Pour aller plus loin :

92 words

Radar Profile

The radar profile shows high scores in information quality and reliability, reflecting the video's rigorous methodology and clear explanations. The quantity of information is also strong, but the technical level is moderate, as the exam is designed for non-science majors. Overall, the video is a well-balanced and informative assessment.

Reliability 8/10

💬 Très positif. Sur les 30 commentaires analysés, la majorité exprime une appréciation pour la démonstration pédagogique et la réflexion sur l'IA dans l'éducation, avec quelques suggestions d'amélioration méthodologique.