One of the biggest problems in computer science education is the following: how do we give implementation assignments when frontier AI models can solve an entire homework from a single prompt written by the student? This article is a collection of approaches that I have encountered through my research while trying to answer this question, together with a somewhat convincing solution that I would like to propose. Note that I specifically consider implementation assignments and attempt to analyze the candidate solutions from the perspective of their nature. However, the same reasoning can be applied to other fields where the main learning outcome of the assignment is problem solving.

The Problem

Let us define the problem in more detail. Most implementation assignments have well-defined instructions and require the student to manually write programs that solve certain problems from the course's domain. Given the current state of AI models, implementation assignments in any course can be solved with a single prompt. Moreover, although it is very easy to figure out whether someone used AI, it is almost impossible to prove it. Therefore, accountable plagiarism detection is not possible.

Another problem is that there is now a significant gap between the models that people can freely access or run and the frontier models provided by AI labs. Even though open-source models exist, the most "intelligent" ones are too expensive to run on commodity hardware. This creates a significant gap among the dishonest students themselves as well. A student with better financial means and access to a Claude Max subscription can complete an assignment in a significantly shorter amount of time than a student using a free-tier account.

Disclaimer: The following phrases "student using LLM" or "dishonest student" in this article refer to the students who completely offload the task of doing the homework to an LLM and submit its output as a solution. Although still not ideal, we are not concerned about students using LLMs to assist with their thinking while doing the homework.

What is evaluation after all?

The purpose of any kind of evaluation within a class is to deliver a message directed at the student. The overall evaluation of a student in a class (a letter grade, a statement of the learning outcomes) is a message to the student's future evaluators. Therefore, the evaluation of the homework should contain a direct message to the student. Before going over possible resolutions of the problem, I think we should define what this message should be. In my opinion, the message is:

"If you are not doing the work yourself, you are not learning."

This is obvious when you simply read it, but the goal is to embed this message into the students' decision mechanism by consistently simulating this experience for them. More importantly, the earlier this message is delivered to the student, the better.

Some students argue: "if AI can do all those things in the homework, what is the point of learning it?" This is not a logical argument, since the homework has a predefined solution, and the goal is to practice arriving at that solution again using what was learned in class. Hence, a homework's value has nothing to do with whether it can be done by someone else; it was always doable by someone else (primarily by the person who prepared it).

Put differently, the objection confuses the value of having the solution, which is zero because the solution already exists, with the value of producing it yourself, which is the entire point of the homework.

Suboptimal Solution Candidates

We now go over the common approaches that I either disagree with or find suboptimal. The solutions below will be evaluated by how well they convey the above message to the student (described under the message).

Do nothing

This was my initial response to the problem. My chain of thought was the following: "offloading the task of doing the homework to AI is the same as paying someone else to do it; therefore, nothing should change." However, this view is wrong on two fronts. First, paying someone else to do your homework is considered unethical by most students, whereas offloading it to AI is not (I think both are equally unethical, but this is not the reality, and it is a discussion point of its own). Second, the evaluation mechanism is now dysfunctional, as described in the problem section, and it is the evaluator's responsibility to put in additional work to make the evaluation functional again.

Message: This sends the exact negation of the message, since a student who offloads the work still receives full credit, and the evaluation itself testifies that the work was unnecessary.

Remove the homework

This can only be done if the learning outcomes of the homework can be transferred to some other part of the course. One can argue that if the homework can be removed very easily without affecting the learning outcomes, then it should have been removed already in the first place (even in the pre-LLM era).

Message: This sends no wrong message, but it sends no message either, since with the homework gone there is no longer any moment in the course where the student can experience that skipping the work means skipping the learning. At least for the material that was supposed to be learned during homework assignments.

Give more difficult homework

One common idea is to expand the scope of the homework so that solving it properly takes many iterations and the management of AI agents. However, this penalizes honest students and forces them not to do the homework on their own. The only exception would be a class whose learning outcomes include agentic coding, agentic problem solving, and the like.

Message: This inverts the message for the honest students. Instead of "do the work yourself", the course now announces that "the work is no longer doable by yourself".

Make the homework optional

This again penalizes the students who would actually like to do the work. Time is a valuable resource in undergraduate education, and assigning a non-trivial homework that takes a couple of hours to solve without giving anything back to the student is not very practical. One practical implementation of this approach could be to give detailed feedback to the students who did the homework voluntarily; that feedback itself could compensate for the absence of a grade.

Message: This demotes the message to a private opinion, since it reaches only the students who already believed the message in the first place and are willing to pay for it in unrewarded time.

Oral exam

This is a valid solution, but it really depends on the course. It is only possible with a small cohort or a very large teaching crew (or by holding the oral exams in groups). Personally, I also think all oral exams should be recorded in some way to prevent objections. Overall, it places an unreasonable amount of work on the teaching crew.

Message: This conveys the message well but through a channel whose cost grows with the cohort, so the message can only ever be delivered to small audiences or with a substantial cost on teaching labor.

Test the homework in class

This is the most sound approach, and it is being adopted in many places. Say one of the homework assignments contributes 10% of the grade. This grade is split into two parts (it could be 5+5 or 2+8, depending on the homework). The first half is given exactly as the homework would have been given in the pre-LLM era. Then, immediately after the homework deadline, there is an in-class exam (probably a quiz) on the homework topics. Notice that this exam is neither part of the midterm nor of the final; there is a dedicated but small exam evaluating the learning outcomes of each homework. For implementation homework, this exam can consist of the student's own submitted implementation printed on paper, with the teacher asking the students to explain their own implementation. However, this pattern of examination is prone to students memorizing the exam pattern very quickly, since they can let the AI do the homework and then memorize the solution, again with AI, without doing much work. Also, constructing a personalized exam (even if the exam generation is automatic) comes with the burden of having a seating plan in the exam in order to distribute the papers properly. Still, this option might work in many classes.

Message: This states the message well if the exam is designed well (more on that below).

My Proposal: Make the Effortless Use of AI Pointless

This approach is very similar to testing the homework in class, but it has some specific details that are designed to make it more effective. I will explain the methodology through a sample scenario; the homework task is oversimplified in order to demonstrate the methodology.

Before AI

Homework assignment. Implement quicksort. Benchmark it on your computer and give running times for several input sizes. Compare your experimental results to what you would expect from the complexity of the algorithm.

Duration. 7 days.

Submission. On Moodle.

Evaluation. Based on the implementation and the experiments.

Grade contribution. 10%.

After AI

Homework assignment. Implement quicksort. Benchmark it on your computer and give running times for several input sizes. Compare your experimental results to what you would expect from the complexity of the algorithm.

Duration. 10 days.

Submission. On Moodle.

Notice that the duration has increased to 10 days. That is because we are going to release a reference solution to the homework on day 7 (yes, you read that right). Here is what we do: on day 7 (the normal duration within which we would expect an honest student to finish the homework), we release a reference solution that we wrote, and we also ask two frontier models (e.g., Claude Fable, GPT 5.6) to generate reference solutions. We upload these three solution documents to Moodle. Students are free to execute one of the reference solutions on their computers and submit the homework using them. The figure below summarizes the timeline.

Two side effects of the day-7 release should be acknowledged. First, after day 7, a student who genuinely needs more than 7 days loses the chance to produce the solution on their own. I consider this an acceptable cost, since day 7 is exactly the deadline we would have set in the pre-LLM era. Second, the released solutions will accumulate publicly over the years. This would have been a serious concern in the pre-LLM era, but today it changes nothing, since every frontier model already "leaks" the solution to anyone who asks.

Timeline of the proposed homework format

Timeline of the proposed homework format. The assignment is released on day 0. On day 7 (the time by which an honest student is expected to have finished), the instructor's reference solution and two frontier-model solutions are published on Moodle. The submission deadline is on day 10 (worth 2%), and it is immediately followed by an in-class quiz (worth 8%).

Evaluation. Submission on Moodle and an in-class exam right after the homework deadline.

Grade contribution. Moodle submission 2%, quiz 8%. The homework assignment technically still has a 10% weight.

This can be called a classwork or a quiz, and it can be held during or after class hours depending on the size of the course. The delicate point is to design this exam in such a way that an honest student who did the homework needs no preparation for it and will likely get 100%, while a dishonest student cannot easily obtain a high score. Designing these exams of course requires more work from the teaching team. However, as I said, it is our responsibility to adjust the evaluation mechanism when the current one is broken, and it is natural that this requires more work.

At this point, you may object that the memorization pattern I criticized in the test-in-class approach applies here as well, since a dishonest student can let the AI do the homework and then study the released solutions before the quiz. This is where the design of the quiz becomes critical. If the quiz fails to distinguish these two students, this experiment might fail. Specifically, getting a high score in the quiz should require more effort (in studying the solutions, or overall) from a student who skipped the homework than from a student who plainly did the homework. Under this condition, doing the homework honestly should be the cheapest path to a high score, and the message we want to give is enforced in practice. The work cannot be avoided, and a student who skips it ends up doing more of it.

I can hear you asking two questions:

  1. "Releasing the homework solution before the deadline is bizarre." You need to realize that, in the post-LLM era, the homework solution is available from day 1 to everyone with access to a frontier model.

  2. "What does this accomplish over testing the homework in class?" It has the additional property of making LLM use pointless. It conveys the message that we would like to give in a very strong manner. More specifically, it says: "The purpose of the homework is not to somehow get access to its solution; we can simply give it away. The real learning happens only when you are actually doing the work that produces the solution."

Conclusion

The core of my proposal is easy to state. In the post-LLM era, the solution to an implementation homework is effectively public from day 1, whether we like it or not, so we should stop pretending otherwise and release it ourselves. Once the solution is officially free, the only thing left worth grading is the one thing that cannot be copied, namely the understanding that comes from doing the work.

I do not claim this is the final answer. The whole construction stands or falls with the quiz. An honest student should be able to walk in with no preparation and get full marks, while a student who skipped the homework should need more effort to reach the same score than the homework would have cost them in the first place. Whether such quizzes can be designed consistently, course after course, is an empirical question and I intend to treat it as one, starting with my own courses.