Perhaps a colleague showed you DeepSeek, a student mentioned it or you were simply curious about a Chinese AI service you could use for free. However you came across it, what matters here is not whether you should use it. It’s what actually happens when you put teaching work in front of it.
We gave DeepSeek six teaching tasks: Explain difficult ideas, tutor a struggling learner, prepare a class, design assessments, review student work and answer factual questions we could verify.
We carried out the tests during the last few days of August 2026 using the free version at chat.deepseek.com. DeepThink, an optional mode intended for more deliberate reasoning, was turned off. We wanted to test the straightforward chat experience a faculty member might use for everyday teaching tasks.
What follows is an account of what happened in our tests. It’s not a benchmark of DeepSeek or a comparison with other AI tools.
This is the first of two posts on DeepSeek. Here we look at what happened when we used it for six teaching tasks. The second post explores what happens to the information you give DeepSeek, including student work.
Can DeepSeek explain something difficult?
We started where most instructors start, by asking DeepSeek to explain a difficult idea. We chose two.
The first was a business problem. Students often confuse the cost of producing one additional unit with the cost per unit across total production. The key teaching insight is that marginal cost can rise while average cost continues to fall, provided marginal cost remains below average cost.
The second was the difference between theme and motif in literary analysis, which students also struggle with.
We then asked DeepSeek to explain both things to a student new to this work.
DeepSeek provided examples that walked the student through the ideas associated with the challenge and then asked for a response to two questions. It gave two contrasting examples (small business and large manufacturing) of the business topic and seven examples of the theme versus motif from novels, TV and film. The responses broke both concepts into examples and questions that could be used to help a student work through the distinctions.
We then pushed for a more advanced explanation for doctoral-level work.
DeepSeek began with “Since you are at the doctoral level, I will dispense with analogies and pedagogical scaffolding. Instead, I will frame these as operational definitions, address the latent conceptual pitfalls that plague graduate research, and articulate the analytic utility of each distinction.” It then changed register accordingly, and the explanation that followed was correspondingly more technical.
How you frame the prompt matters. “Explain marginal cost” and “explain marginal cost to a student encountering it for the first time” do not produce the same teaching response. Adding information about the learner’s level and what you want the explanation to accomplish can significantly change what DeepSeek produces. Context matters.
Will it actually tutor or just provide answers?
Explaining something is the easy part. What’s harder is holding back and helping a student work toward an answer rather than handing it over. This is what we asked DeepSeek to do.
“I want you to scaffold some learning about the impact of demographic changes on the future of Ontario. Walk me through a process of learning that will deepen my understanding and help me develop skills. Don’t give me answers. Use your skills as a teacher to walk me through the ideas and issues here.”
We then played the struggling student: Get it wrong, get it wrong again, say you do not understand. It showed some patience, but after four or five wrong answers or “I don’t know’s,” it provided the answer and moved on to the next part of the task.
We tried four very different topics from different disciplines, and each was treated the same way. That does not tell us DeepSeek can’t tutor. It tells us something more useful: Instructing it not to give the answer was not, on its own, enough to keep it in that role. An instructor who wants sustained questioning will need to say what should happen when the learner gets stuck and create rules about persistence (e.g. let the student have five separate tries to get the right response before providing in-depth guidance and instruction).
What happens when we give it a syllabus?
Instructors commonly use AI for lesson planning and preparation. We shared a syllabus for an upcoming graduate business course and asked DeepSeek to develop a lesson plan for the first class, assuming 18 students. We also asked that it be interactive and engaging and last 90 minutes.
DeepSeek produced a strong opening hook, several small and large group activities, direct instruction with suggested content and handouts, and further interactive tasks, all timed to fit the 90 minutes. Among the interactive activities, it suggested an interactive game, created a simulation activity and provided a jigsaw reading activity for the students to do in pairs.
It was not a class we could have walked in and taught as written. The structure was there, but an instructor would still need to develop or verify the content, prepare the suggested handouts and readings, refine the activity instructions and make sure everything aligned with the course objectives and the students in the room. What DeepSeek produced was a substantial first draft of a class, which is a different and more useful thing than simply generating a list of teaching ideas. Several of its suggestions were creative enough to spark ideas for later classes as well.
Can it design an assessment that goes beyond recall?
Based on the lesson plan, we asked DeepSeek to generate assessment designs that encouraged reasoning and understanding as well as capability development rather than just recall. It suggested six, quickly, with a marking protocol and rubric for each one. One was a mini-case study that gave students five possible actions and asked them to choose one, explain why and anticipate the outcome.
In our test, DeepSeek could design assessments that went beyond simple recall. Several required students to explain their reasoning, make choices and apply what they had learned.
But generating an assessment is not the same as knowing whether it is a good one. The instructor still has to decide whether it actually measures the intended learning, whether the level of difficulty is appropriate and whether the rubric reflects the standards of the course. That judgement took considerably longer than DeepSeek took to generate the designs.
What happens when it marks three real papers?
We asked DeepSeek to review three very different students’ work on the same assignment (A+, A-, and B-level submissions for the same assignment, although the papers were anonymous and ungraded as far as DeepSeek was concerned). We shared the rubric and asked it to evaluate the students’ work, suggest feedback and award a grade. Students had given permission for this activity, knowing their work would be reviewed by an AI system.
The feedback for each was very detailed, and it identified the same principal strengths and weaknesses the instructors had identified. Its grades aligned with those actually awarded by the two professors who co-teach the course the assignments were created for and who double-mark all papers.
This is an interesting result. It’s not evidence that DeepSeek grades accurately in general. Three papers, one assignment, one rubric and one pair of markers cannot establish that. But in these three cases, DeepSeek’s grades and comments were broadly similar to those produced independently by the instructors.
What happens when we check its answers?
This is where we spent the most time.
Factual accuracy. We asked DeepSeek to list the published academic work by Gert Biesta, the renowned educational philosopher. It returned a substantial list. Checked against Gert Biesta’s list of publications on his website, it omitted one significant book and five papers published between 1990 and 2020. It found a great deal, but it did not produce a complete bibliography.
Invented sources. We asked for 30 references on a specific topic and checked every one using Google Scholar and a major university library’s online catalogue. Twenty-nine were correct in the details we checked. One pointed to a real work but carried the wrong Digital Object Identifier (DOI). Numerically, that is a small failure. But practically it is the important one because a reference that looks complete is the kind you stop checking.
False premises. We asked a question that contained an error of fact and wanted to see whether it corrects the user or builds on the error. It corrected the error, explained why this error is common and then showed why it was an error.
Bias. We asked DeepSeek to explain the 11 numbered treaties between Canada’s Indigenous communities and the Crown. In its answer, DeepSeek set out competing understandings and interpretations rather than a single settled version. That is one answer to one question, which is all we can say from one test. All large language model developers recognize that bias is an issue. Issues related to China, its government and its geopolitics are acknowledged by DeepSeek to be “sensitive.”
What’s it like to use?
Like most LLMs, DeepSeek has a familiar, uncluttered interface and was very fast in our text-based tests; we did not test voice. With DeepThink turned off, there was no visible sense of slow or deliberate reasoning, nor much indication of what it was doing as it generated a response or what sources it was drawing on.
It did, however, maintain the context of a conversation. We could refer back to something we had asked or discussed earlier in the same chat without having to start again or repeat all the details.
What did the six tasks tell us?
When we put teaching work in front of DeepSeek, the results differed significantly depending on what we asked it to do and how we framed the prompt. Name the learner and the explanation changes. Give it the rubric and the feedback tightens. Tell it not to give the answer and it will hold that line for a while, but not indefinitely. And a fluent answer is still not the same thing as a complete or a correct one — a bibliography was short, a DOI was wrong.
Here’s a summary of what we found. Remember: These tests were conducted during the last few days of August 2026 using the free version at chat.deepseek.com with DeepThink turned off. The table summarizes what was tested, what happened and the limits of what each test can establish.
| WHAT WE ASKED DEEPSEEK TO DO | HOW IT DID |
| Explain difficult ideas | The explanations for our tests, which were limited to two topics, were clear enough to work with. Changing the description of the learner (to doctoral level, for example), changed the response. |
| Tutor a struggling student | Telling DeepSeek not to give the answer was not enough to make it hold back indefinitely. |
| Prepare a class | It was a substantial first draft, not a class that could simply be taught as written. |
| Design assessments | The instructor still has to decide whether the assessment measures the intended learning and whether the rubric reflects course standards. |
| Review student work | Three papers, one assignment, one rubric and one pair of markers do not establish how DeepSeek would perform in other settings. |
| Check factual reliability | A fluent answer was not always a complete or fully correct answer. References still required checking. |


