Friday, January 11, 2008

Chinese Cyberwarfare (ISN)

Rachel Kesselman, a Mercyhurst grad student, just published an interesting piece of analysis on Chinese cyberwarfare with the International Relations and Security Network (ISN). (Way to go, Rachel!)

Part 8 -- Confidence Is Not the Only Issue (The Revolution Begins On Page Five: The Changing Nature Of The NIE And Its Implications For Intelligence)

Part 1 -- Welcome To The Revolution
Part 2 -- Some History
Part 3 -- The Revolution Begins
Part 4 -- Page Five In Detail
Part 5 -- Enough Exposition, Let's Get Down To It...
Part 6 -- Digging Deeper
Part 7 -- Looking At The Fine Print

Part 8 -- Confidence Is Not the Only Issue

Some 29% of the sentences in the Iran National Intelligence Estimate (NIE) do contain Words of Estimative Probability (WEPs), however. As the chart below shows, this is pretty much in line with other NIEs. The chart outlines the number of uses of a particular word in an estimative sense in each of the eight NIEs I examined. Again, I only looked at the words in the Key Judgments (not in any of the prefatory matter or in any of the full text or appendices). The column on the far right shows the percent of the time a particular WEP showed up in NIEs generally. In other words, "probably" was used in 33 sentences and there were 263 sentences total in the 7 NIEs examined, so it showed up about 13% of the time. I am also well aware that such a simple review is fraught with difficulty given the complexity of the English language but, since I am only looking for broad trends, I believe that such a review is an appropriate method for analyzing the way in which these estimates were written and the way in which they are changing.





In fact, the Iran NIE is well within the range of other NIEs with respect to percent of sentences containing WEPs. Furthermore, the Iran NIE does not use any “unauthorized” WEPs. That is to say, only WEPs specifically listed on the Explanation of Estimative Language (EEL) page are used in the Iran NIE. This was not the case in previous NIEs which used (though not often) statements that were undefined at a minimum and misleading at their worst. Consider the use of “most likely” in the August 2007 update to “Prospects for Iraq’s Stability”:

  • We judge such initiatives are most likely to succeed in predominantly Sunni Arab areas, where the presence of AQI elements has been significant, tribal networks and identities are strong, the local government is weak, sectarian conflict is low, and the ISF tolerate Sunni initiatives, as illustrated by Al Anbar Province.

“Most likely” could mean many things in this context since there is no baseline probability with which to compare it. The initiatives referenced in the report could be likely to succeed or unlikely to succeed; the reader cannot know from the text. All we can know is that they are "most likely" to succeed in the predominantly Sunni areas. Other formulations, such as “much less likely” and “increasingly likely”, suffer from the same problem. “Not likely” is the only place where I am clearly quibbling as it is obviously synonymous with unlikely. I just think it is silly to state that the authors intend to use “unlikely” on page 5 (the EEL page) and then ignore that and use “not likely” in the text. If the two are truly synonymous then use the one you said you were going to use. If they aren’t synonymous, then explain the difference. You can’t have it both ways.

Beyond the mere use of WEPs, there also appears to be an issue with which WEPs predominate. Again, there is a strong pattern – the clear preference over the last 6 public NIEs for the use of the word “probably”. In fact 73% of authorized and 62% of all WEPs used in the last six NIEs are “probably”. It is also interesting to note that the only non-millennial NIE examined, the 1990 Yugo NIE did not use “probably” at all (whether this pattern holds and whether this was a good thing, I will leave to other researchers).

If the analysts involved in these estimates genuinely believe that all these events are “probable” and not somewhat more or less likely then there is little to discuss. The extreme overuse of the term suggests other explanations, however. "Probably" is arguably one of the broadest WEPs in terms of meaning (see Figure 1 in the paper linked here). Fairly clearly it means that the odds are above even chance but it seems open to interpretation from there.

Thus, analysts could be using "probably" as an analytic safe haven. Relatively certain that the odds are above 50% but unwilling to be more aggressive and use a phrase such as “highly likely” or “virtually certain” and unaware or unable to use expressions of confidence to appropriately nuance these more aggressive terms, these analysts default to “probably”. Since the NIE is a consensus estimate combining input from all 16 intelligence agencies, it is also possible that "probably" was the one word upon which everyone could agree; that it represents, essentially, a compromise position. Either way, such a move is “safe” in terms of getting the answer broadly correct but hurts the decisionmaker who, in the end, must take action and allocate resources. If analysts are more certain than they are willing to put in writing, the decisionmaker is deprived of the analysts’ best judgment and will arguably make less informed decisions.

(Note: The statistical analogy to the issue described above is the classic problem of calibration versus discrimination. For additional insights into this issue I refer you to Phillip Tetlock’s book Expert Political Judgment or to this site)

Monday: Part 9 -- Waffle Words And Intel-Speak

Thursday, January 10, 2008

Iran-Navy Clash New Footage (IRGC via PressTV)

According to PressTV ("the first Iranian international news network, broadcasting in English on a round-the-clock basis"), the Iranian Revolutionary Guard Corps not only claims the incident with the US Navy was fake but is now providing its own video to prove it. You can view the video and report here or can click on this link to open the video in your windows media player (Many thanks to M. for the lead!)

Talking To The Enemy (RAND)

Dalia Dassa Kaye at RAND has recently (late 2007) written a very insightful report (click here for full text) on the value of so-called "track two diplomacy" -- unofficial contacts between ordinary people on different sides of an issue -- in the Middle East and South Asia. I am familiar with some of Dalia's previous work and she has always been worth the read. While the full report has much to ponder, here are some highlights from the Summary (Boldface is mine):

  • "While official diplomatic communications are the obvious way for adversaries to talk, unofficial policy discourse, or track two diplomacy, is an increasingly important part of the changing international security landscape."
  • "The experiences of the Middle East and South Asia suggest that track two regional security dialogues rarely lead to dramatic policy shifts or the resolution of long-standing conflicts. But they have played a significant role in shaping the views, attitudes, and knowledge of elites, both civilian and military, and in some instances have begun to affect security policy. However, any notable influence on policy from such efforts is likely to be long-term, due to the nature of the activity and the constraints of carrying out such discussions in regions vastly different from the West."
  • "Track two dialogues on regional security are less about producing diplomatic breakthroughs than socializing an influential group of elites to think in cooperative ways."
  • "Track two dialogues typically involve moderate and pragmatic voices that have the potential to wield positive influence in volatile environments, and the stakes are high."
  • "This study identifies three conceptual stages that define the evolution of track two dialogues, although in practice these stages are not necessarily sequential: socialization, filtering, and policy adjustment."
    • "During socialization, outside experts, often from Western governments or nongovernmental institutions, organize forums to share security concepts and lessons based on experiences from their own regions."
    • "Filtering involves widening the constituency favoring regional cooperation beyond a select number of policy elites involved in track two, through the media, parliament, NGOs, education systems, and citizen interest groups. In practice, this stage has often been the weak link in track two dialogues, as there has been inconsistent translation of the ideas developed in regional security dialogues to groups outside the socialized circle of elites."
    • "The final stage is the transmission of the ideas fostered in dialogues to tangible shifts in security policy, such as altered military and security doctrines or new regional arms control regimes or political agreements. Track two has not led to such extensive shifts in security policy, although there are examples of track two work influencing official thinking and a variety of security initiatives and activities, particularly in South Asia."
  • "Track two dialogues in the Middle East have affected growing numbers of regional elites."
    • "What have these dialogues achieved over the years? Their socialization function has succeeded in shaping a core and not-insignificant number of security elites across the region to begin thinking and speaking with a common vocabulary."
    • "That said, the filtering of track two concepts has by and large failed to penetrate significant groups outside the dialogue process."
  • "As in the Middle East, South Asia experienced a growth in track two dialogues in the 1990s, and many of these efforts continue today."
    • "The direct impact of South Asian dialogues on official policy has been limited, although not entirely absent."
    • "A number of confidence-building measures (CBMs) initially discussed in track two forums are now being officially implemented between India and Pakistan, such as the ballistic missile flight test notification agreement, military exercise notifications and constraint measures along international borders, and Kashmir-related CBMs."
    • "South Asian dialogues have also succeeded in changing mindsets among participants toward more cooperative postures and have had some success in building a constituency supportive of South Asian cooperation, including in challenging areas such as nuclear confidence building and new approaches to Kashmir."
    • "Filtering is also apparent from the emergence of a variety of regional policy centers focused on issues that are being discussed in track two venues."
  • "Still, track two groups in both regions have made considerable progress in socialization. Thousands of military and civilian elites have discussed and engaged in cooperative security exercises. Expertise and knowledge of basic arms control concepts were limited in both regions before the 1990s. Now, because of track two dialogues, there are large communities of well-connected individuals familiar with such concepts. Knowledge of complex arms control and regional security concepts and operational confidence-building activity is now solidly rooted in both regions."

Part 7: Looking At The Fine Print (The Revolution Begins On Page Five: The Changing Nature Of The NIE And Its Implications For Intelligence)

Part 1 -- Welcome To The Revolution
Part 2 -- Some History
Part 3 -- The Revolution Begins
Part 4 -- Page Five In Detail
Part 5 -- Enough Exposition, Let's Get Down To It...
Part 6 -- Digging Deeper

Part 7 -- Looking At The Fine Print

Let’s take a look at an example from the Iran National Intelligence Estimate (NIE) and see if we can figure out what is going on here.

  • We judge with high confidence that Iran will not be technically capable of producing and reprocessing enough plutonium for a weapon before about 2015.

What is missing in this example and many other statements in the NIE, of course, is an estimate of likelihood (or Word Of Estimative Probability (WEP), if you prefer). The estimate does not say “…will not likely be technically capable …” Instead, the verb phrase “will not be technically capable” implies certainty about a future event – which is, by definition, uncertain.

Even in cases where the event happened in the past but the information regarding the event contains inconsistencies or uncertainties (in other words the event is not definitively factual) such as this statement (also from the Iran NIE), “We assess with high confidence that until fall 2003, Iranian military entities were working under government direction to develop nuclear weapons”, it seems inappropriate to not use estimative language in conjunction with a statement of confidence.

In other words, if the Intelligence Community (IC) knew for certain that Iranian military entities were working under government direction to develop nuclear weapons then they should not be indicating a probabilistic statement by saying “we assess”. If, on the other hand, they are not entirely certain, then they should not say “Iranian military entities were working” but rather “it is virtually certain that Iranian military entities were working” or whatever the analysts believe is the appropriate estimate of likelihood. Mixing the formulations makes the definitions laid out in the Explanation of Estimative Language (EEL) page -- the "page five" in the title -- meaningless.

This is problematic for several additional reasons. First, it is bound to be confusing to the reader. Having carefully explained that “We judge” is an indicator of estimation but then phrasing the statement in terms of certainty makes the attentive reader wonder what the IC really means; is this statement a fact or an assessment? It could be read both ways. Second, another graduate student with whom I have worked, Mike Lyden, has done some very interesting research comparing NIE estimative statements against historical fact (his thesis that contains the research is available currently only through inter-library loan). Across the last 40 years, estimates that used WEPs tend to be about 75% accurate. Statements that use words of certainty hover around 50% accuracy (the sampling size was large enough that this difference was statistically significant to several decimal places as I recall). Mike speculates that this difference may be tied up with psychological notions of confidence (explained later) but whatever the reason, the evidence is pretty compelling – the Intelligence Community makes better estimates when it does not use words of certainty.

Another possibility, of course, is that I have got it all wrong; that I have mischaracterized what the IC intended to do when they defined “confidence” the way they did. Indeed, there are several other ways that the word "confidence" could be interpreted that would work in this sentence.

First, confidence could refer to psychological confidence or the way the analyst "feels" about the assessment. Psychologists have long known that the more information you get the more confident you feel in your assessment of a situation. Up to a point, this increasing confidence is warranted. Fairly quickly, however, your mind forms a more or less rigid conceptual model of the problem you are facing so that your mind takes each new fact and tends to either force it into the existing model or discard it as irrelevant. The net effect of this is that, while you feel increasingly confident, your chances of being correct stay about the same. Psychologists call this Overconfidence Bias and it is generally considered a bad thing in analysis. Moreover, it is well known within intelligence circles, having been covered extensively by Richards Heuer in his classic, Psychology Of Intelligence Analysis. It is, therefore, unlikely to be what the IC means when it talks about confidence on the EEL page.

Second, confidence is often used as a synonym for likelihood as in “I am highly confident that New England will win the Super Bowl.” While this works in casual speech, this certainly makes no sense in the context of this NIE. The EEL page defines an entirely different way of ascribing levels of likelihood to its assessments and specifically states that the level of confidence language applies “to our assessments (italics mine).” To use confidence as a synonym for likelihood would be tantamount to the IC saying one thing and doing another which, well, they have already done. I don’t, however, think they would be that silly again. For the same reason, the introductory phrases, “we assess”, “we judge” and “we estimate” can’t be considered to be expressions of likelihood either.

Third, and likely most closely related to what the IC means, is a statistical notion of confidence, commonly expressed as a margin of error. The form of the statement is quite familiar to most of us: “Candidate X leads in the polls, 61 to 39% (plus or minus 3 percent).” This means (typically at the 95% confidence level – yet another statistical term) that Candidate X’s true lead could be as low as 58% or as high as 64%. This form certainly seems to mirror the form examined in Part 4. High confidence under this interpretation would mean that the margin of error is low, that the true probability hovers near the estimate made by the authors of the estimate. The problem here comes in the way the IC has actually used confidence in these phrases. If they mean it to be interpreted statistically it makes no sense to then say something that would be functionally equivalent to “…plus or minus 3 percent, Iran will not be technically capable…”. This kind of statement and others like it only make sense when associated with a probability or, in the case of the NIE, an Estimate Of Likelihood.

This, in turn, brings me back to the more general notion of analytic confidence that I discussed in Part 4. Certainly the IC does not want to convey numerical certainty and has said so (at least in early forms of the EEL page) but this idea of analytic confidence seems similar to the idea of statistical confidence. By using words (not numbers) that express likelihood and then using words (not numbers) to express its confidence in an expression of likelihood, the IC’s implied definition of analytic confidence would resonate with, but not mirror, what many people already generally understand, i.e. the statistical notion of confidence. Just as with statistical notions of confidence, however, this idea of analytic confidence only makes sense if there is an expression of likelihood to go with it.

Which leaves me with a problem. I don’t know what the IC means when they talk about confidence. The EEL page implies they intend to use it one way. Then they do something entirely different in the text and none of the possible variations in meaning makes any sense. They do it so many times that I can’t ascribe it to accident.

I am just an average Joe. The first alternative is that I just don’t understand. I am prepared to admit that. I would suggest, however, that the current form of the EEL page needs to be changed so that it is clearer. I guarantee that if I cannot understand what it means, there are many more average Joes that are struggling with it (or just ignoring it) as well.

The second alternative – and one that is a bit more unnerving – is that the IC does not know what it means when it says high, moderate or low confidence. Perhaps sometimes they are using it to describe how they feel about their position, sometimes they may be using it as a synonym for an estimate and sometimes they may mean it more statistically, leaving it up to the reader to figure out which it is from the context.

Tomorrow: Part 8 -- Confidence Is Not the Only Issue