The Socrates Project: Assessment Redesign in the Age of AI

Abstract: Facing a crisis of academic dishonesty due to generative AI, the US Army’s Command and General Staff College initiated the “Socrates Project.” This project introduces a novel assessment method where students must use an artificial intelligence agent, Socrates, to generate their papers through a guided Socratic dialogue. By evaluating the student’s raw cognitive process during the conversation rather than the final written artifact – such as an essay or a report – this approach effectively deters cheating and provides deeper insights into students’ true understanding and learning capabilities.
Higher education institutions across the country face a crisis as more students use generative artificial intelligence (GenAI) to cheat or take short cuts which result in cognitive decline and reduced academic rigor. A study from May 2026 found that 26% of students use GenAI daily with 9% of students using GenAI to cheat. While the number of students using GenAI to cheat might continue to grow, the study notes that “These patterns call for discipline-specific assessment reform, not blanket bans or universal detection regimes.” While higher education and professional military education institutions grapple with this dilemma, the US Army’s Command and General Staff College (CGSC) implemented a novel assessment method to deter students from using GenAI to cheat while also creating entirely new forms of assessments. This resulted in the “Socrates Project.”
The Socrates Project serves as a template for how higher education institutions might employ AI agents to deter and defeat the use of GenAI as a tool for students to cheat or bypass important cognitive steps required for learning. The project addresses the problem by requiring all students to an AI agent – Socrates — to generate their final paper.
Socrates was born out of the results of the Athena case study at CGSC. Athena is an AI agent created by soldier-developers that can grade student papers and, “perform the grading task with a high degree of consistency and alignment to human graders.” The true end goal of Athena was to validate a hypothesis: Could an AI agent assess a student paper or exam against a rubric as accurately and consistently as a human grader? The evidence in the Athena case study validated that hypothesis, allowing CGSC to redesign assessments for an AI-era.
If an AI agent can assess an artifact as well as a human grader, then an entire new world of possibilities for assessments open to educators. What the Socrates Project sought to test was, “For assessments, can we mandate students use Socrates, an AI agent, and have that agent assess their conversation and their visible cognitive process as the means of assessment itself?” Instructors assess artifacts because educators have limited means to assess a student’s raw cognitive process outside of oral boards. Such oral boards are immensely valuable for assessments, but are logistically burdensome, time intensive, inconsistent across instructors, and difficult to organize and execute at scale. Socrates is designed to directly address these problems.
The Design Philosophy of the Socrates Agent
Socrates is an AI Agent created in Gemini on GenAI.mil with a robust set of instructions and training in the Socratic method of instruction. The instructions for Socrates were general enough to accept any rubric and — using the Socratic method — ask a series of questions assessing against that rubric’s criteria and the appropriate level of learning per Bloom’s Taxonomy, an educational framework used to structure, measure, and scale the depth of a student’s critical thinking. Once Socrates assesses the student has reached a minimum of a B grade, it informs the student and offers to generate their artifact for them, using only the information in the answers that the student provided throughout the process of Socratic questioning and dialogue. This approach flips the role of GenAI tools to function as a developmental editor rather than a ghostwriter.
Socrates is designed to allow students to continue their assessment and keep their conversation going if they wish to obtain a higher grade. This approach is informed by the transparency and consistency of the Army Fitness Test (AFT). For example, soldiers taking the AFT know exactly how many repetitions or what time standard they need to meet to achieve the score they desire. On the other hand, academic assessments sometimes approach assessment without a clearly defined rubric; it is as if students were assessed on push-ups with no knowledge of the exact repetitions required to maximize their score. Moreso, when it comes to consistency, each faculty member grades somewhat subjectively and arbitrarily. Some students might receive higher or lower scores simply due to the grading style of the faculty member grading them.
The core design hypothesis of Socrates is this: can we create enough transparency to give a student an idea of their grade throughout the assessment process so that they can demonstrate what the assessment requires of them? This design allows two things to occur at once. Firstly, faculty can gain an assessment of a student’s knowledge and learning style. This also provides feedback on the student’s observable cognitive process and level of motivation. Secondly, it provides students with confidence that assessments are fair, balanced, and without faculty bias. Many soldiers will push themselves further when others cheer them on and count their repetitions out loud. These soldiers perform better with these factors than they would have without any knowledge of their score. Socrates replicates this competitive boost in an academic assessment and assesses its effect on student motivation.

In addition to transparent grading, Socrates provides the student with instant feedback on strengths and areas for improvement while explaining the thinking process behind the grade. Often, feedback arrives after the student begins their next block of instruction or assignment, too late to be useful. Additionally, Socrates generates an assessment for the faculty member that details the depth and complexity of the student’s intellectual work as well as strengths and weaknesses. This assessment allows faculty members to mentor and coach each student in a way that best meets their educational needs. For each of these assignments, Socrates produced an information paper artifact using the student’s answers to its questions with a hyperlink to the conversation embedded in the artifact as well as student feedback and a faculty assessment report. The student submitted the information paper artifact with the embedded hyperlink to that conversation and the faculty assessment report to the faculty member.
Why submit the faculty assessment report to the faculty member? If their conversation went poorly, would the student submit their work or restart Socrates until they got it right? In the latter case, the student would still learn and improve. If a student submits a negative faculty assessment to their instructor, it says a lot about the student and informs the faculty member beyond what they might otherwise gather. This approach allows faculty to gain more insight into the student than ever before.
Methodology
This iteration of the Socrates Project involved over120 students across two cohorts of the AI Basic Course at CGSC. This course is a new elective where a student builds an AI agent to address problems identified by their gaining unit. Students use Socrates to create a two-page information paper based on their AI agent. While this information paper is the final artifact, it includes a hyperlink to the precipitating conversation with Socrates. This conversation is locked and auditable by faculty; it is the substance of this conversation from which the information paper artifact is derived. Instructors grade students on a unique rubric as assessing a conversation is fundamentally different than grading an information paper.
The rubric for this assignment has six categories to assess both the student’s conversation with Socrates and the information paper artifact. Categories 1–5 evaluate the quality of the student’s reasoning and engagement during the Socratic conversation. Category 6 evaluates the information paper artifact, reflecting the quality of the conversation itself. Each rubric category has four score ranges ranging from exceptional to unsatisfactory and a definition of each grade level. Additionally, each category has its own weight, with the information paper artifact, category 6 weighted at 10% and the other 90% derived from the categories 1-5.
Cohort 1: Validating Socrates
For cohort 1, Socrates only provided an estimated grade threshold, such as a “B” or an “A” letter grade. It did not provide an exact percentage; the faculty member was ultimately responsible for assigning a percentage grade. Once the student reached a B threshold, Socrates offered to create their information paper artifact. However, the student could continue if they wanted to pursue an “A.” If a student could not reach a B threshold, Socrates refused to create their information paper artifact and would encourage them to continue their conversation. Once a student was ready, Socrates produced the information paper artifact, faculty assessment, and student feedback then submitted their assignment.
To ensure its validity, the instructor manually read each student’s conversation with Socrates and graded each exam using the rubric. On average, each conversation was 15 pages. The insights gained on each student were far better than those from a two-page information paper. Below are a few insights observed during this manual review process:
Socrates Validated
After reviewing and grading over 700 pages of conversation text, Socrates maintained consistency with the rubric. Its feedback for both the student and faculty was sound and based on relevant observations from the conversation. This, plus the work from the Athena case study, allowed Socrates to incorporate a percentage grade for cohort 2.
Raw Cognition
Students threw proper grammar out the window in these conversations. In an age of grammatically correct GenAI writing, this was a breath of fresh air and a way to see real thinking happening during an assessment. If we want to deter the use of GenAI to cheat or mask true intellectual capability, what better way than by requiring raw thought as the means of assessment? How long have academics ignored creative thinkers and eccentric intellectualism in favor of strict grammatical correctness? For once, the student’s true intellectual capability shined through in their raw cognition during the process of creating an artifact, a process that defied assessment until now.
Human-Machine Teaming
For this assignment, we authorized some students to use another GenAI tool. It was fascinating to witness human and machine cognition interact as they conversed with Socrates. Students had to input questions from Socrates, then digest, copy, and paste them into another GenAI tool, then paste the refined output back into Socrates. Grammatically incorrect writing appeared occasionally as students poked and prodded, figuring things out and learning in the process. Ultimately, we could easily identify more even if a student used GenAI tools than without them. Perfect, polished grammar is easily identifiable from unpolished conversational text.
Cohort 2: Improving Socrates
Once we validated that Socrates could consistently assess student conversations, we refined and improved Socrates for cohort 2. We created two sub-agents, an evaluator and a report compiler. The evaluator sub-agent is objective, neutral, and highly analytical, providing precise, data-driven percentage scores based strictly on the provided rubric and the student’s conversation with Socrates. The report compiler activates at the end of a session after the information paper generates and the student receives a final grade from the evaluator sub-agent. It assembles the faculty assessment report using inputs provided by Socrates and the evaluator sub-agent and does not interact with the student beyond the completed report.
This creates an autonomous agentic assessment ecosystem: a student converses with Socrates, an evaluator assesses, and a report compiler observes. This ensures that the grade given by the sub-agent evaluator is not tied to Socrates or the student tricking Socrates into providing a higher grade. Additionally, the report compiler sub-agent reports any observed misconduct in the faculty assessment report.
The grade provided by the evaluator sub-agent is the source of the most interesting interaction in the new Socrates design. When a student asks for their grade, Socrates relinquishes control to the evaluator sub-agent that informs the student of their grade at that moment. The student can finish or continue the assignment as they see fit or pursue a higher grade. When a student is ready to end their assessment, the evaluator sub-agent provides the student with a final grade and asks the student to concur or disagree with their grade. If a student disagrees, they are asked to provide a statement that is included in the faculty assessment report for consideration. This is not for the student to negotiate their grade, only a request for a human-in-the-loop. It also allows us to observe how many students concur or disagree with the evaluator sub-agent’s assessment. Lastly, the report compiler includes the number of times a student asked for their grade in the faculty assessment report.
For this second cohort, we spot-checked each conversation in lieu of a manual review. Our review of the artifacts and faculty assessment reports, we made the following observations:
Grade Transparency and Competitiveness
The students in this cohort asked for their grade an average of three times throughout the exam. In the first cohort, all students stopped once Socrates informed them that they had reached the A grade threshold, and only 20 students achieved 100% on the assessment we manually graded. In cohort 2 however, most students kept asking for their grade and continued until they reached a score of 95%. A total of 26 students continued until they reached 100%. Some students wrote comments like “I want to keep going!” or “I am going to go all the way until I get 100%!” They were actively pushing themselves, competing for the 100% as if this was a competition; it appeared as if they were motivated to beat Socrates. How many of our students could earn a higher grade if they knew what was expected of them? We would prefer to assess what a student can accomplish when pushed instead of what they display for assessment.
Insight into Human Behavior
There were observable behavior patterns that provided much more insight into students than simply assessing an artifact. We learned which students were driven to get a 100%, which kept pushing until they got to 95%, and which stopped when they got an 84%. Students’ raw conversations with Socrates were particularly insightful. One student answered the questions well and clearly knew the material but just did not go deep enough. Why was this student the only one who stopped at an 84%? This could give us an indicator to check on our students, understand their personal life, and support them as faculty. In this instance, the student was overwhelmed with movers and preparing to begin their next assignment. We would not have this insight if we were just grading an analogue information paper artifact.
Results
The students from both cohorts requested more Socratic assessments be a part of CGSC. Moreso, most students found value in their conversations with Socrates and improved their own AI agents, the original intent of the assignment. It is rare for such an assessment to result in applied learning.
With grading, we believe there is significant value in the autonomous grading ecosystem of Socrates in cohort 2. Ultimately, the mean and median grades between the two cohorts were nearly identical. However, there was a difference in the number of students earning 100% on their assessment. In cohort 2, many students continued their conversations with Socrates just to pursue a 100% on the exam. When we look at their conversations, students declare this indirectly in their justification to keep their conversation going. The following table provides a more detailed breakdown of the results from each cohort:
| Cohort | Mean | Median | 100% Grades | Min Grade | Std Dev |
| Cohort 1 | 97.78% | 98.00% | 30.8% (20) | 93.00% | 1.79% |
| Cohort 2 | 97.60% | 98.00% | 45.6% (26) | 84.00% | 2.91% |
Conclusion
The Socrates Project has profound implications for higher education broadly and professional military education specifically. For higher education, the Socrates Project shows that we can leverage GenAI to create agents that conduct new and novel forms of assessments. Educators must innovate and adapt to the AI era by creating new forms of assessments for a student’s raw cognitive process. The final artifact carries less value the same way the answer to a math problem now carries less weight after the advent of calculators. The cognitive process, or showing your work, must become the standard of assessment. This will deter the use of GenAI to cheat and mask intellectual capability.
For professional military education, the Socrates Project has the potential to uncover the officers and leaders that traditional means of assessment might miss. How many top-graduates fail on the battlefield while poorly performing students become notable combat leaders? Could Socrates have identified creative and innovative thinkers like General George Patton who struggled in traditional assessments focused on assessing polished artifacts? Can the military leverage GenAI to assess the raw cognitive processes of its officers and leaders to best assign talent in the future? There are practical implications beyond academic research or behavioral science. For educators and military leaders, this topic requires further study and exploration. The Socrates Project teaches us that traditional forms of assessment will die in the AI-era. The most powerful means of educational and cognitive assessment will be born from those ashes.
The views here are those of the author and do not represent the opinions or positions of the CGSC, the U.S. Army, the Department of Defense, or any part of the U.S. government.