AI as a Staff Instrument
Lessons from a Shadow Staff
By MAJ Alex Noll and Richard A. McConnell, D.M.
| Field Artillery, 2026 E-Edition
Read Time: < 12 mins
U.S. Army 1st Lt. Mahdi Al-Husseini, assigned to the 25th Combat Aviation Brigade, pitches his presentation on AI pilot performance feedback. Al-Husseini is one of seven Soldiers who is taking part in Dragon’s Lair 5. The program was established in October 2020 to help increase innovation across the XVIII Airborne Corps. (Photo by Sgt. Marygian Barnes)
“AI is – as those of us building it like to joke – ‘what computers can’t do.’ Once they can, it’s just software.” – Mustafa Suleyman, The Coming Wave, p. 951
Introduction
Our Information Collection Practicum article2 made the case that artificial intelligence (AI) works best when a human designs the process around it. In that platform test, an AI G2 persona helped a student staff group process intelligence collection requirements, connect Master Scenario Events List injects to named areas of interest, and produce intelligence assessments. The following platform test asked: Could the same process-based approach replicate an entire staff? During the W400 and W500 practicum sequence, the answer was “sort of.” The Shadow Staff did not replace human planners. Instead, the AI was best used as a partner whose usefulness depended on the humans around it, what they structured, what they calibrated and what experience they brought. The platform test reveals that effective AI integration in military education depends less on model capability and more on the human processes built around it.
A tempting claim is that AI could replicate an entire staff, especially when a model can quickly produce estimates, matrices, course of action statements or wargame products. Perhaps someday it will, but the more useful claim is narrower: AI can simulate selected staff functions when the human team gives it doctrine and context, sets the constraints, and builds a workflow that exposes errors.3 In that sense, the Shadow Staff served as a foil to the human staff: It made human doctrine, calibration, authority and judgment stand out more clearly. In practice, the platform test produced three roles for AI throughout this process: (1) Blue Staff planner, (2) Red Staff adversary and (3) White Cell adjudication aid. Each role helped students see the planning problem from a different angle, and each revealed where human authority still must own the decision.
Role 1: Blue Staff Planner
The first role was the closest extension of the IC Practicum platform test. In W419, the AI environment operated as a Blue Staff planning partner for a division-level defense. The system used nine specialized agents4 organized around staff functions and warfighting functions with a chief-of-staff orchestration layer coordinating the effort. The value was not raw-output volume. It was structure: procedural gates5 that enforced doctrinal sequence, checked feasibility and prevented a COA from moving forward when a key staff function had not yet made a required determination. For example, a sustainment feasibility gate kept the planning honest by forcing logistics to be addressed before the product could be treated as complete.
The result was faster, but bounded, staff production. However, speed was not the central lesson. Across two sessions, the Blue Staff architecture produced more than 40 planning products (such as receipt of mission, initial staff estimates, mission analysis slides, recommended priority intelligence requirements) and allowed multiple warfighting-function estimates to develop and be refined in parallel. That compression matters in a school environment, where the AI-assisted learning helped the staff develop deeper coherence across time-constrained planning. Students could see how intelligence, fires, sustainment, protection, signal and maneuver interacted rather than treating each product as an isolated assignment.
The system also had a hard boundary: It could structure commander's intent and surface its implications, but it could not originate intent or resolve the commander's guidance on its own. Using the Shadow Staff effectively required the students to think through how the process shifts slightly to better facilitate commander decisions. For example, a tactics, techniques, and procedures (TTP) the class developed was to have the AI operate a half- to full-step ahead with anticipatory drafting, while ensuring the next step still fell in line with the most current commander's guidance. This required an involved AI operator who stayed tied in with the commander and the staff. As the students drew concept sketches on a whiteboard, the AI operator was capturing each team’s rationale, assumptions, potential decisions and open questions. This allowed the AI to draft a 15- to 20-slide course-of-action (COA) development brief, with COA statements, recommended tasks and purposes by phase, decision support matrices, sync matrices and risk assessments.
This workflow provided two critical aids. First, during COA development, more students were involved in the planning and less concerned about ensuring the briefing products were being built in a briefable format simultaneously. Second, it gave students more time to think through the actual concept. By the time the sketch was complete, the slides were generated. While slides were never perfect on the first iteration, when the sketches were complete, each COA team would open its respective slides and make quick changes to its respective portions of the brief.
Role 2: Red Staff Adversary
The second role required a more significant rebuild. W500 was not a simple reskin of the Blue Staff system. The Red Staff was designed as an adversary planner built around Olvanan-inspired planning logic,6 with 12 agents7 using different staff designations and different priorities. The design forced the system to behave differently from a U.S. staff. Electronic warfare (EW) and information operations (IO) were treated as condition-setting efforts prior to fires and maneuver. EW received its own agent; deception and systems destruction provided the means of creating exploitation windows, and statecraft appeared as a distinct planning concern to influence both the Blue Forces (BLUEFOR) and the local populace.
This architectural change did two things. It produced more realistic opposing forces (OPFOR), and it gave students a working aid for understanding OPFOR doctrine.8 The Red Staff used a systems frame that put the enemy into subsystems and looked for nodes whose disruption would create cascading effects. While the outputs were similar to the BLUEFOR, what proved more important was the intellectual friction it created. BLUEFOR planners had to contend with an opponent that did not think like they did, and OPFOR planners had a sparring partner that helped them better simulate an Olvanan staff. This provided both teams with a learning environment better than just using blue tactics as the opposing force.
Lastly, the Red Staff demonstrated why AI-supported adversary work still needs human judgment. Some products were structurally correct but culturally shallow. This was highlighted through the integration of the Information Advantage (IA) Scholars during the planning process. The IA Scholars9 provided better focus for the information operations (IO) agent and better shaped the information space with refined expertise. The agents are only as good as their instructions, and having the IA team embedded in our planning helped refine both the understanding of the audience and of the information environment.
Role 3: White Cell Adjudication Aid
The third role was the most concrete test of performance. During the W500 wargame, the White Cell used AI to help process player-submitted land power10 products, identify engagements, apply combat results tables and prepare adjudication recommendations. This role was different from planning. It required more mechanical accuracy, consistent naming and more careful context engineering for each turn. The first calibration run showed the problem clearly. The AI properly parsed only about 2/3 of the data accurately. Creative player-unit names broke the matching logic. That initial result served as a reminder that trust requires calibration through human involvement in any AI process.
The team then invested time to make the tool more accurate and useful. The unit aliases were added; a skill was created11 that matched them to the official names, numeric designators, slash-separated references and partial matches (e.g., 2-34AR, 2-34, 2/34). After calibration, the system matched over 90% of cells across 10 tabs in an Excel workbook. The remaining differences were not all errors. Some were dice outcomes, instructor adjudication or judgment calls. The White Cell workflow did not remove humans from adjudication. It separated the mechanical number crunching and data transfer from the judgment work and made the handoff between the two explicit.
A final element added to the White Cell adjudication workflow was a confidence checker. AI can help process volume but only if the process shows where confidence ends. For this approach, we used an idea from CW4 Jacob Land from Fort Sill. Through his example, we learned how he used a script to adjudicate exercises and produce operations summaries and intelligence summaries (OPSUMs and INTSUMs) for a tabletop exercise (TTX) they were running. Using that idea as a baseline, we created two skill files12 : one to generate reports and another to identify low-confidence areas across the workbook. The color-coded triage approach distinguished routine mechanical adjudication from events requiring instructor attention. This allowed the White Cell to focus less on building the fastest spreadsheet and more on producing a credible fight that supported learning. In future iterations, a calibrated AI aid could free instructors from some number-crunching tasks while making the moments requiring human judgment more visible.
What the Shadow Staff Did Not Solve
The platform test also clarified three limits worth noting: intent, calibration and pedagogical friction. First, AI cannot originate commander’s intent. It can help structure the implications of intent, test whether a plan follows that intent, and identify gaps between stated priorities and staff products, but without clear intent, it’s just moving faster in the wrong direction. The best way to establish this context is for the Shadow Staff to receive guidance directly from the commander or instructor, which can get complicated if the guidance is not written or recorded.
Second, the Shadow Staff ’s ability to produce doctrinally correct outputs does not mean it has the refined contextual experience of humans. Human planners were needed to provide that experience, to spot check the outputs and to recalibrate the agents when necessary. An excellent example was when the signal agent produced insufficient products for the group (i.e., they took longer to edit than just make it themselves). One of the signal officers sat down at the Shadow Staff computer and updated the agent’s instructions. After about 20 minutes of recalibrating and editing, the agent was able to build products that only needed minor tweaking.
Lastly, AI will try to solve every problem you put in front of it. In a school setting, that can quietly erase pedagogical context unless humans build it back into the workflow. A student exercise is not only a tactical problem but also a learning environment. If the instructor wants to introduce friction or ambiguity for students to think over, this “instructor intent” should be built into the tool. If the AI smooths away those moments, it can make the exercise cleaner while making it less educational. The same caution applies to adversary and cultural nuance: Plausible outputs still need human judgment to test whether an enemy would actually think, message or act that way.
Recent work using AI with Mind Genomics13 reinforces this limitation. Mind Genomics breaks complex issues into smaller elements, recombines them into short vignettes and tests how people respond. If we pair that with AI personas, then a staff can determine how different framings might produce different reactions before committing to one. However, the output remains a hypothesis, not a validated truth. For the Shadow Staff, the value is not predicting what commanders, adversaries or populations will do; it is identifying which assumptions deserve a second look.
Implications for PME and the Force
The Shadow Staff platform test suggests that professional military education should treat AI more as a staff instrument. Whether students can use a model to draft a paragraph is settled—they can. The harder question is how we teach them to design a process that makes AI useful without making it authoritative. That requires doctrinal understanding, structured inputs, validation measures, calibration and a habit of human review.
For students, the payoff is repetition. They can see their plan through friendly, enemy, and adjudicator perspectives inside one exercise, get faster feedback, and confront an adversary that is structurally different from themselves. For faculty, it is relief from number crunching and swivel-chair work and more time spent on coaching judgment and decision making. For the force, the lesson is that useful AI integration will not come from prompts. It will come from the human-AI workflows built around them.
Conclusion
The IC Practicum article14 began with a claim that a better process15 matters more than the technology. The Shadow Staff platform test reinforces that claim at a larger scale. AI did not replace the staff. It helped simulate selected staff functions, exposed inconsistencies, accelerated some mechanical tasks and created a more demanding learning environment. Its value came from the human investment around the tool—what got structured, what got validated, and what was kept for human judgment. The Shadow Staff did not automate the staff; it forced the staff to explain its judgment when the AI recommended a different course.
Endnotes
1. Mustafa Suleyman and Michael Bhaskar, The Coming Wave (Crown, 2023).
2. Alex Noll and Richard McConnell, “Designing AI Integration: A Process Based Approach,” Army.mil, March 1, 2026, https://www.lineofdeparture.army.mil/journals/field-artillery/field-artillery-archive/field-artillery-2026-e-edition/designing-ai-integration/.
3. Forrest A. Woolley, “The Case for Engineering an AI Partner for Intellectual Honesty in the National Security Ecosystem,” Small Wars Journal by Arizona State University, April 30, 2026, https://smallwarsjournal.com/2026/04/30/the-case-for-engineering/.
4. Claude Code Docs, “Create Custom Subagents - Claude Code Docs,” Claude.com (Claude Code Docs, 2026), https://code.claude.com/docs/en/sub-agents.
5. Claude Code Docs, “Hooks Reference - Claude Code Docs,” Claude.com (Claude Code Docs, 2026), https://code.claude.com/docs/en/hooks.
6. Headquarters Department of the Army, “File:TC 7-100.2 - Opposing Force Tactics (December 2011).Pdf - Wikimedia Commons,” Wikimedia.org, December 2011, https://commons.wikimedia.org/wiki/file:tc_7-100.2_-_opposing_force_tactics_(december_2011).pdf.
7. Claude Code Docs, “Create Custom Subagents - Claude Code Docs,” Claude.com (Claude Code Docs, 2026), https://code.claude.com/docs/en/sub-agents.
8. Headquarters Department of the Army, “File:TC 7-100.2 - Opposing Force Tactics (December 2011).Pdf - Wikimedia Commons,” Wikimedia.org, December 2011, https://commons.wikimedia.org/wiki/file:tc_7-100.2_-_opposing_force_tactics_(december_2011).pdf.
9. The Information Advantage Scholars program has a historical partnership during the capstone exercise at CGSC with teaching team 6. While the information space is less of a focus during the planning sessions at CGSC, this partnership provided Team 6 with exposure to how their actions influence competing narratives across the battlespace.
10. Land Power is a map-based analog wargaming simulation used at CGSC to help teams fight plans against one another after planning sessions are complete.
11. Claude Code Docs, “Codebase Explorer,” Claude.com, 2026, https://code.claude.com/docs/en/skills.
12. Claude Code Docs, “Codebase Explorer.”
13. Howard Moskowitz, “Researchopenworld.com,” Researchopenworld.com, March 3, 2026, https://researchopenworld.com/israel-hamas-and-strategy-focused-experiments-using-ai-and-mind-genomics/.
14. Noll and McConnell, “Designing AI Integration.”
15. David De Cremer and Garry Kasparov, “AI Should Augment Human Intelligence, Not Replace It,” Harvard Business Review, March 18, 2021, Kasparov, March 19, 2021, https://www.kasparov.com/ai-should-augment-human-intelligence-not-replace-it-harvard-business-review-march-18-2021/.
Authors
Alex Noll, MBA, is a Major in the Infantry and currently a student in Staff Group 6A at the Command and General Staff College at Fort Leavenworth, Kansas. He has served in both airborne and light assignments and received his MBA from the Raymond A. Mason School of Business at William & Mary while serving in the Army Futures and Concepts Center.
Richard A. McConnell, DM, is a retired Army Lieutenant Colonel and a professor in the Department of Army Tactics U.S. Army Command and General Staff College at Fort Leavenworth, Kansas. He served as the principal investigator for the summer 2022 creativity study dedicated to exploring ways to improve creativity among students. The creativity study research report was published in the 2023 Association for Business Simulations and Experiential Learning (ABSEL) Conference proceedings. He received his DM in organizational leadership from the University of Phoenix and has published several articles on wargaming, exceptional information, creativity and ethics-related topics.