Beyond the Admiralty Code: Rethinking Information Evaluation in Civil Affairs

Introduction
Across the US Army’s modernization efforts, increasingly central to decision-making at every echelon. The vision of a “digitally-enabled, data-driven” force, however, rests on the precondition that the data feeding the Army’s systems are reliable. This article examines that precondition through the framework that Civil Affairs (CA) uses to evaluate civil data as part of civil knowledge integration (CKI) and proposes an updated approach.
Senior leaders have called data “the new ammunition,” and data centricity has become a focal point of Army transformation. Models built on poor data will produce flawed analyses. Additionally, the volume of available civil data has simply outpaced humans’ capacity to evaluate it manually. As technology increasingly augments reporting, the branch needs scalable methods to assign relative to, and filter, the data it relies on.
Whether analysis is conducted by traditional means or emerging tools, the quality of the data drives the quality of the insight. This matters acutely in complex and dynamic civil environments. CA provides commanders with expertise on the civil component of operational environment (OE)—the civil knowledge that shapes battlefield decision-making. Army Techniques Publication (ATP) 3-57.50: Civil Knowledge Integration (2024) defines data as “unprocessed observations from the environment,” information as “data that has been processed to provide context,” and knowledge as “information analyzed to provide operational implications.”
According to Field Manual (FM) 3-57: Civil Affairs Operations (2021), CKI is “the actions taken to analyze, evaluate, and organize collected civil information for operational relevance and informing the warfighting function. The resulting civil knowledge is integrated with other knowledge about the OE to create shared understanding among commanders, unified action partners, international organizations, and civilian partners.” That knowledge is incorporated through the Army’s integrating processes of intelligence preparation of the battlefield (IPB), information collection, targeting, risk management, and knowledge management, which together enable the commander’s understanding of the OE and develop the common operational picture.
Because civil knowledge is integrated into these processes, the quality of the underlying data can significantly impact mission outcomes. For example, a brigade combat team (BCT) operating among large numbers of dislocated civilians (DCs) may need them directed away from military movement routes toward designated receiving locations. The staff may assess that mobile communications can reach the population and that nongovernmental organizations (NGOs) can provide essential services.
However, that assessment may rely on coverage maps, organizational directories, and other sources that are incomplete or outdated. Cellular service may not reach portions of the affected population, while listed NGOs may no longer operate in the area, work in the required sectors, possess sufficient capacity, or be willing to coordinate with military forces. If civilians cannot be reached or expected services are unavailable, they may continue to use routes required for military movement, delaying maneuver and sustainment forces and placing the mission at risk. The underlying data may appear authoritative, but if their reliability and currency have not been properly evaluated, they can produce an inaccurate civil picture for the commander.
Current doctrine provides a framework for evaluating civil data at collection but offers no means of accounting for temporal currency. Borrowed from Second World War-era human intelligence (HUMINT), it was also built for a very different information environment without digital sources, automated processing, or the risks that come with AI integration.
The Current Framework: A Brief Historical Overview
ATP 3-57.50 (2024) places civil data evaluation criteria in step two of CKI, “collect civil data,” before formal analysis and evaluation in step four. At collection, the framework applies two A-F ratings: one for “reliability” and another for “accuracy” (see Table 1 and Table 3).
This framework is not native to CA, but a lightly modified Admiralty Code first developed by British Naval Intelligence during the Second World War. The US Army adopted it in the 1951 Field Manual (FM) 30-5: Combat Intelligence to evaluate the credibility and accuracy of human informants. It now underpins HUMINT standards for source reliability and information credibility. Table 2 and Table 4 show the current matrix from FM 2-22.3: Human Intelligence Collector Operations.
Placed side by side, the CA and HUMINT frameworks are essentially the same instrument. CA renamed the second axis “accuracy,” changed numbers to letters, and adapted some vocabulary for civil data, but most descriptors and the underlying logic remain unchanged.
The Admiralty Code extends beyond US and British practice. NATO standardized it across the alliance (also referred to as the NATO Intelligence Source Evaluation Code), where it remains a widely used shorthand for source evaluation in Western intelligence. The 2003 NATO Standardization Agreement (STANAG) 2511 codified the framework, later updated in the 2016 NATO AJP-2.1 (see Table 5).
Table 1: CA CKI Reliability Ratings (ATP 3-57.50)
| A | Reliable | There is no doubt the civil data is authentic, trustworthy, or complete. There has been a history of complete reliability. The civil data is within adherence to known professional standards and a verification process. |
| B | Usually reliable | There is minor doubt that the civil data is authentic, trustworthy, or complete. There is a history of valid information most of the time. The civil data may not be adhering to professionally accepted standards. |
| C | Fairly reliable | There is doubt that the civil data is authentic, trustworthy, or complete, but information provided in the past has been valid. |
| D | Not usually reliable | There is significant doubt that the civil data is authentic, trustworthy, or complete, but information provided in the past has been valid. |
| E | Unreliable | The civil data is lacking in authenticity, trustworthiness, or completeness. Other civil data indicates this civil data is invalid information. |
| F | Cannot be judged | No basis exists for evaluating the reliability of the source. |
Table 2: Evaluation of Source Reliability (FM 2-22.3)
| A | Reliable | No doubt of authenticity, trustworthiness, or competency; has a history of complete reliability. |
| B | Usually reliable | Minor doubt about authenticity, trustworthiness, or competency; has a history of valid information most of the time. |
| C | Fairly reliable | Doubt of authenticity, trustworthiness, or competency but has provided valid information in the past. |
| D | Not usually reliable | Significant doubt about authenticity, trustworthiness, or competency but has provided valid information in the past. |
| E | Unreliable | Lacking in authenticity, trustworthiness, and competency; history of invalid information. |
| F | Cannot be judged | No basis exists for evaluating the reliability of the source. |
Table 3: CA CKI Accuracy Ratings (ATP 3-57.50)
| A | Confirmed | The civil information is confirmed by other independent sources. The information is logical in itself. The information is consistent with other information on the subject. |
| B | Probably true | The information is not confirmed. The information is reasonably logical in itself. The information agrees with some other information on the subject. |
| C | Possibly true | The information is not confirmed. The information is possible but not logical. There is no other information on the subject. |
| D | Doubtfully true | The information is not confirmed. The information provided is possible but not logical. There is no other information on the subject. |
| E | Improbable | The accuracy of the information is not confirmed. The information is not logical in itself. The information is contradicted by other information on the subject. |
| F | Cannot be judged | No basis exists for evaluating the validity of the information. |
Table 4: Evaluation of Information Content (FM 2-22.3)
| 1 | Confirmed | Confirmed by other independent sources; logical in itself; Consistent with other information on the subject. |
| 2 | Probably true | Not confirmed; logical in itself; consistent with other information on the subject. |
| 3 | Possibly true | Not confirmed; reasonably logical in itself; agrees with some other information on the subject. |
| 4 | Doubtfully true | Not confirmed; possible but not logical; no other information on the subject. |
| 5 | Improbable | Not confirmed; not logical in itself; contradicted by other information on the subject. |
| 6 | Cannot be judged | No basis exists for evaluating the validity of the information. |
Table 5: NATO Source Reliability and Information Credibility Scales (NATO AJP-2.1)
| Reliability of the collection capacity | Credibility of the information | ||
| A | Completely reliable | 1 | Completely credible |
| B | Usually reliable | 2 | Probably true |
| C | Fairly reliable | 3 | Possibly true |
| D | Not usually reliable | 4 | Doubtful |
| E | Unreliable | 5 | Improbable |
| F | Reliability cannot be judged | 6 | Truth cannot be judged |
A Critique of the Current CA Framework
As currently applied, the framework is not well-suited for civil data evaluation for three interconnected reasons: corroboration, currency, and consistency of interpretation.
Corroboration
CA defines itself as separate from the intelligence apparatus (ATP 3-57.50 states plainly that “CKI is not an intelligence activity”); however, its civil data evaluation framework is a HUMINT instrument built to judge human informants. In HUMINT, the second axis measures credibility through corroboration and whether a report is “logical in itself.” CA renamed this axis “accuracy,” but retained criteria that measure credibility rather than factual accuracy.
“Credibility” earns a separate axis in HUMINT because a human source can relay a single report that proves more or less reliable than its record. But this is not equally applicable to civil data because a source generates its data through its own methodology, so a dataset’s credibility is largely inseparable from how it was produced. This is also why reliability can be judged at the point of collection while credibility cannot be judged at that stage independently. A source’s methodology, independence, and scope are disclosed in its own documentation, which is accessible when the data are gathered. Establishing credibility, by contrast, means evaluating the content against other sources, which is a process that requires reaching beyond the source and one that CKI assigns to its later “analysis and evaluation” step. For civil data, then, the credibility axis is either redundant with reliability or premature at the point of collection.
The corroboration standard is also difficult to meet because civil data often have only a single authoritative source. This is common with official government reporting, where an independent equivalent may not exist. An axis topping out at “confirmed by other independent sources” therefore sets a standard unreachable in many cases. For example, an operator working from the sole host-nation report on paved-road percentages cannot confirm that figure short of inspecting the country’s road network.
Currency
Information currency—the degree to which information remains representative of present conditions as a function of its age, the OE, and the rate at which the underlying metric changes—is a related and arguably more consequential problem. Built to assess a human informant, the current framework makes no allowance for information age. For civil data, that omission is significant because the age of a data point is often central to its utility. A figure can come from a reliable source and be correct when collected, but a decade later its relevance depends on the metric’s volatility and the OE. This is especially true in large-scale combat operations (LSCO), where conditions change constantly.
Consistency
The third problem is subjective terminology, a long-standing critique of the Admiralty Code and the field writ large. ’s 1964 “Words of Estimative Probability” found that intelligence analysts attached markedly different probabilities to terms such as “serious possibility.” A 1975 Army Training and Doctrine Command (TRADOC) study similarly found wide variation in how intelligence officers interpreted the Admiralty Code, and the authors recommended the development of a new scale. More recent scholarship has identified the same semantic problem: without defined values, one analyst may interpret “usually reliable” differently from another. These critiques apply equally to CA.
The CA adaptation illustrates the problem directly. In the accuracy ratings (see Table 3), the descriptors for “possibly true” and “doubtfully true” are identical except for one word: “the information is possible” versus “the information provided is possible.” An operator has no criteria with which to distinguish a C from a D rating.
A Proposed Updated Framework
These critiques point to three requirements for an updated framework: source reliability criteria tied to observable characteristics rather than subjective labels; temporal currency; and recognition that information degrades at different rates depending on the OE. The revisions below address these requirements while preserving the simplicity that makes the current framework accessible in practice.
Revised Axis 1: Source Reliability
The source reliability axis retains the A-F scale and classifications from ATP 3-57.50, but grounds each level in criteria operators can assess directly (see Table 6). Predominantly favorable criteria indicate A or B, mixed criteria C or D, and predominantly unfavorable criteria E. Where the criteria cannot be assessed, F applies. The four criteria are:
- Transparency of the source’s methodology.
- Validation or assurance (e.g., an independent audit or peer review).
- Independence from external influence or bias that could shape reporting.
- A clearly defined scope of coverage, with documented reasons for any gaps, including access or collection limitations.
These descriptors apply across civil data sources, including official host-nation statistical bureaus, international organizations (IOs), NGOs, academic and research institutions, and private-sector entities. Each carries distinct challenges: host-nation statistics may be politically influenced, even among allies; NGO reporting may reflect advocacy bias; and private-sector data may offer limited methodological transparency. Where criteria conflict, such as a source with documented methodology but no external validation, operators weigh the source’s overall profile and assign the rating that best reflects its demonstrated reliability. Grounding each level in observable criteria reduces the subjectivity that has long been a weakness of the Admiralty Code.
Table 6: Proposed Updated CKI Reliability Ratings
| Code | Label | Descriptor |
| A | Reliable | Methodology is documented and publicly available. Validation or assurance is provided. of external influence or bias. Scope is defined, data gaps are explained (e.g., access or collection limits). |
| B | Usually reliable | Methodology is documented with minor gaps. Validation or assurance is provided, but limited or partial. Substantially independent with limited potential for external influence or bias. Scope is defined, with some data gaps unexplained or incomplete. |
| C | Fairly reliable | Methodology is documented but materially incomplete. No validation or assurance. Identifiable external influence or bias. Scope is partially defined, with data gaps noted but unexplained. |
| D | Not usually reliable | Methodology is absent or undisclosed. No validation or assurance. Significant external influence or bias. Scope is undefined, and data gaps are undisclosed. |
| E | Unreliable | Methodology is absent. No validation or assurance. Not independent, with documented evidence of fabrication or systematic misrepresentation. Scope is undefined, and data gaps are undisclosed. |
| F | Cannot be judged | Insufficient basis to evaluate against these criteria. Source is new or unreviewed, and no material is available on which to assess it. |
Revised Axis 2: Information Currency
The accuracy axis is difficult to apply at collection because it requires independent confirmation that often does not exist and asks operators to judge whether information is “logical in itself,” frequently without the subject-matter expertise to do so. In practice, this often yields “F: cannot be judged.” Accuracy remains important, but operators can seldom establish it at collection. The proposed framework relocates that judgment to the later “analyze and evaluate” step, restoring the sequence CKI already prescribes.
In place of “accuracy,” the proposed framework substitutes “information currency,” which captures how old the data are relative to the pace of change in the OE. Unlike accuracy, currency is more observable, measurable, and often more relevant at the point of use. Ratings are based on collection dates rather than publication dates (although the two can coincide) and run on an A-F scale.
However, not all civil data age at the same rate. Land areas and administrative boundaries can remain stable for decades; population and infrastructure tend to change over years; and during conflict or natural disasters, displacement and economic indicators can shift within days, weeks, or months. These differences affect how currency should be applied, so the framework classifies metric volatility at three levels:
- Static: Metrics that do not meaningfully degrade with age. Confidence should follow source reliability alone.
- Gradual: The default for most civil metrics, and the rate of change against which the currency table’s were calibrated.
- Dynamic: Metrics that are event-driven or rapidly changing. For these, the currency table is a baseline rather than a final answer, especially in crisis and LSCO, where a metric rated “current” may already be operationally stale. Dynamic metrics require the greatest operator judgment.
Because volatility governs how the currency rating is applied, it should be assessed first. Once volatility is classified, the currency rating (see Table 7) is applied across competition, crisis, and LSCO, with time ranges calibrated to each environment’s pace of change.
In competition, civil systems are generally stable and collection follows established schedules, allowing a relatively long collection-to-use window. In crisis, systems are stressed, collection may be disrupted, and conditions change faster. In LSCO, displacement, infrastructure damage, and governance continuity tend to be dynamic, so the time range compresses most sharply and requires the greatest operator judgment.
The time ranges reflect general patterns in civil data collection cycles and are practical starting points rather than empirically derived thresholds. They can be refined as the framework is applied and validated in practice. When a rating is in doubt, the less favorable rating should apply, as it is better to understate currency than to overstate it.
Table 7: Proposed CKI Information Currency Ratings
| Code | Label | Competition | Crisis* | LSCO* |
| A | Current | ≤3 years | ≤6 months | ≤3 months |
| B | Mostly current | >3-5 years | >6-12 months | >3-6 months |
| C | Marginally current | >5-7 years | >1-2 years | >6-12 months |
| D | Dated | >7-10 years | >2-4 years | >1-3 years |
| E | Significantly dated | >10-12 years | >4-5 years | >3-5 years |
| F | Historical baseline or cannot be judged | >12 years or unknown | >5 years or unknown | >5 years or unknown |
*Crisis and LSCO thresholds are most sensitive to metric type and require operator judgment. The time ranges for these environments should therefore be viewed as non-binding starting points.
Combined Notations
The two axes combine into a single notation while each remains visible: a low reliability rating signals concern with the source, while a low currency rating signals concern with how well the data reflect current conditions. Both are lettered, retaining CA’s existing convention and distinguishing the notation from the Admiralty Code’s letter-number pairing. It also leaves a third, numeric position open for a credibility rating assigned at the analysis and evaluation step (e.g., AA2), should the branch consider that addition in the future.
Two examples illustrate the notation. A national statistics bureau report with documented methodology and external audit, collected within the past year in competition, would rate AA. Separately, a population dataset that is more than five years old, drawn from a crisis-environment source with documented political interference, would rate DF. Both the source and data age raise concerns, so the figure should be treated as indicative only.
The notation also functions as structured metadata that downstream systems, including AI-assisted analytical platforms, can read to weight or filter civil data when generating insights and reports.
Implications for Automated Analysis
Automated systems can retrieve information that matches a query, but without a proper framework they cannot evaluate source reliability or information currency. AI-driven models without explicit guidance will reproduce manual source-quality problems at far greater volume. Because the proposed ratings are compact and rule-based, automated systems can use them as metadata to filter sources below a threshold, weight reliable and current data, and display ratings alongside generated products for the end user to review.
Conclusion
As the Army moves toward a data-driven, AI-enabled force and CA grows more dependent on large volumes of digital civil data, the branch needs sound evaluation protocols built on observable criteria. The proposed revisions ground source reliability in four criteria, replace the accuracy axis with a temporally calibrated currency rating, and add metric volatility classifications for varying rates of change across environments. They also produce a notation that automated systems can read and act on directly. Together, these changes are intended to help the branch evaluate civil information with greater rigor and meet the standard of reliability that the broader force increasingly demands.