ChatGTP-4's attempt at parts 3 & 4 of the Pre-Exam 2023

We also let ChatGPT provide answers to parts 3 & 4 of the Pre-Exam.

Some practical difficulties were that only plain text could be entered, which means that in part 3, the figures couldn't be provided, and that the underlining when indicating amendments was not present.

If one would consider our answers (see seperate posts, which are not necessarily correct!!) as the intended solution and applied the Pre-Exam's marking scheme (all correct: 5 points, 1 wrong: 3 points, etc), ChatGTP achieved 16/50 points which is well below the required passing grade of 35 (normalized from 70 for the entire exam to 35 for parts 3 & 4) and rather a score associated with mere guessing.

Generally speaking, ChatGPT seems to be able to answer basic novelty and scope questions reasonably well but fails at more specific topics such as ranges. Also, ChatGPT appearantly disagrees with the EPO's relatively strict approach to Art. 123(2) and considers most of the amendments as supported. Perhaps this reveals ChatGTP's US-origin?

See below for ChatGPT's answers and short reasoning (red marked where ChatGPT's answer deviated from ours):

ChatGTP-4 finds the pre-exam legal part quite challenging.

From the ChatGTP-4 paper, we understand that its ability to pass exams has been greatly increased. For example, "on a simulated bar exam, GPT-4 achieves a score that falls in the top 10% of test takers. This contrasts with GPT-3.5, which scores in the bottom 10%.", see [1]. Naturally, we were curious to see how the AI would do on the EQE. The pre-exam legal part seems to be the most accessible for an AI, so this is the part we tried. 

We did two runs, one with a short prompt, and one with a long prompt. Both prompts explain the exam's requirements and ask for step-by-step reasoning. The long prompt also contains a question from last year's exam with an example of the required step-by-step legal reasoning. The prompt had to be very clear the program always needs to answer True or False, otherwise you get a lot of answers explaining why the available information is insufficient to make a decision.