Your concerns are not uncommon: Skogvoll and Odden (2026) found that university physics teachers' concerns about generative AI is as a threat to genuine learning and assessment.
The 'ability' of Large Language Models (LLMs) to 'do' physics is an area of focus for some Physics Education Research groups. The results are mixed, and each 'new' version shows 'improvements'. The research focus has been on LLMs' 'performance' on different types of exam (most notably concept inventories, e.g., Polverini & Gregorcic, 2025, but also short answer written exams, e.g., Pimbblett & Morrell, 2024), and the potential usefulness to teachers for e.g., automating grading (see, Physical Review Physics Education Research's Focused Collection in Artificial Intelligence Tools in Physics Teaching and Physics Education Research; but consider that attention to feedback may rely on the mutual recognition of giver and receiver as humans interacting within a disciplinary context; Corbin et al., 2026). Students' actual use has received much less attention.
A student using an LLM could certainly 'pass' a physics degree (Pimbblett & Morrell, 2024), and multimodal LLMs' performance on concept inventories is typically between student and expert levels (Polverini & Gregorcic, 2025). However, there is quite strong dependence on the exact concept inventory (the 'cost' of using a model does not correlate much to performance), but Polverini and Gregorcic (2025) did not discuss whether material very similar to the questions and answers of the concept inventory may be common in the training data of, or otherwise accessible to (e.g., through web searches) later models and might be reflected in the LLM's 'performance'.
Kilde-Westberg et al. (2026) explore how upper secondary school students used ChatGPT to help with lab tasks, finding that they use it to help with conceptual understanding, not practical issues. Perez Linde and Cardenas (2025) meanwhile report that generative AI tools can allow students to use advanced methods at a much earlier stage than in traditional syllabi. However, when it comes to physics 'understanding', 'misconceptions' such as those reported about vertical motion of a thrown teddy bear reported by Gregorcic and Pendrill (2023) may well remain. For example, Dunlap et al. (2025) reported several dissatisfying features of ChatGPT-4 responses to minimally stated problems around expected motion of objects on an inclined plane: answers might be numerically correct, but the process of the problem 'solving' reported by GPT-4 was often unsatisfactory, most notably that assumptions were not stated.
Pimbblett and Morrell (2024) argue that the general availability of generative AI (LLMs) should prompt physics educators to make significant changes to assessment practices (note, laboratory work is one place where LLMs do not do well: a computer cannot manipulate equipment). Reports from the UK indicate that students may not be as keen about using generative AI as instructors presume, and that, with unclear and often contextually inappropriate guidelines, students are developing their own ethical guidelines for what use is appropriate for them (Kernohan, 2026; Dickinson & Marshall, 2026).
So, how worried need you be as a physics educator? Perhaps not too much: students still want to learn, and learning physics is still more effective when discussions are with other humans than with generative AI (Tong et al., 2025). However, it is clear that the availability of generative AI does make some serious reconsideration of what is taught when (more advanced methods may be introduced earlier), the emphasis accorded to physics reasoning (logic), skills (note taking and recording processes), and problem set up (assumptions, diagramming), and how physics students are assessed necessary. In addition, as an instructor, it will may be worthwhile to occasionally see what LLMs 'say' about topics where students are likely to get stuck to check what information students may be getting from them (and whether it is aligned with or contradictory to your expert teaching).
References:
Corbin, T., Tai, J. & Flenady, G. (2026) The pedagogical power of recognition in feedback, Teaching in Higher Education, https://doi.org/10.1080/13562517.2026.2633634
Dickinson, J. & Marshall, M. (2026) Trained to stop learning: How students are experiencing assessment and learning in an age of AI https://wonkhe.com/blogs/trained-to-stop-learning-how-students-are-experiencing-assessment-and-learning-in-an-age-of-ai/ Accessed 17/07/2026; see also the report linked there.
Dunlap, J. C., Sissons, R. and Widenhorn, R. (2025) Descending an inclined plane with a large language model, Physical Review Physics Education Research 21, 010153, https://doi.org/10.1103/PhysRevPhysEducRes.21.010153
Gregorcic, B. and Pendrill, A.-M. (2023) ChatGPT and the frustrated Socrates, Physics Education 58(3) 035021, https://doi.org/10.1088/1361-6552/acc299
Kernohan, D. (2026) Are students really that keen on generative AI? https://wonkhe.com/wonk-corner/are-students-really-that-keen-on-generative-ai/ Accessed, 17/07/2026.
Kilde-Westberg, S., Johansson, A. and Enger, J. (2025) Generative AI as a lab partner: A case study, Physical Review Physics Education Research 21, 020119, https://doi.org/10.1103/ggy1-3kjk
Perez Linde, A. J. and Cardenas, R. E. (2025) Going beyond by using AI in an introductory physics course, European Journal of Physics, 46(5) 055701, https://doi.org/10.1088/1361-6404/adeb15
Pimbblet, K. A. and Morrell, L. J. (2024) Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees, E uropean Journal of Physics, 46(1) 015702, https://doi.org/10.1088/1361-6404/ad9874
Polverini, G. and Gregorcic, B. (2025) Multimodal large language models and physics visual tasks: comparative analysis of performance and costs, European Journal of Physics, 46(5) 055708, https://doi.org/10.1088/1361-6404/ae03f8
Skogvoll, V. and Odden, T. O. (2026) How physics professors use and frame generative AI tools, European Journal of Physics 47(2) 025713 https://doi.org/10.1088/1361-6404/ae55e0
Tong, D., Jin, B., Tao, Y, Ren, H., Islam, A. Y. M. A., and Bao, L. (2025), Exploring the role of human-AI collaboration in solving scientific problems, Physical Review Physics Education Research 21, 010149, https://doi.org/10.1103/PhysRevPhysEducRes.21.010149