Showing posts with label ACH. Show all posts
Showing posts with label ACH. Show all posts

Tuesday, March 25, 2014

Reduce Bias In Analysis: Why Should We Care? (Or: The Effects Of Evidence Weighting On Cognitive Bias And Forecasting Accuracy)

We have done much work in the past on mitigating the effects of cognitive biases in intelligence analysis, as have others. 

(For some of our work, see Biases With Friends, Strawman, Reduce Bias In Analysis By Using A Second Language or Your New Favorite Analytic Methodology: Structured Role Playing.)
(For the work of others, see (as if this weren't obvious) The Psychology of Intelligence Analysis or Expert Political Judgment or IARPA's SIRIUS program.)

This post, however, is indicative of where we think cognitive bias research should go (and in our case, is going) in the future. 

Bottomline: Reducing bias in intelligence analysis is not enough and may not be important at all. 

What analysts should focus on is forecasting accuracy. In fact, our current research suggests that a less biased forecast is not necessarily a more accurate forecast.  More importantly, if indeed bias does not correlate with forecasting accuracy, why should we care about mitigating its effects?

In a recent experiment with 115 intel students, I investigated a mechanism that I think operates at the root of the cognitive bias polemic: Evidence weighting. 

Having surveyed the cognitive bias literature, key phrases began to stand out such as:
A positive-test strategy (Ed. Note: we are talking about confirmation bias here) is "the tendency to give greater weight to information that is supportive of existing beliefs" (Nickerson 1998). In this way, confirmation bias not only appears in the process of searching for evidence, but in the weighting and diagnosticity we assign to that evidence once located.
The research of Cheikes et al. (2004) and Tolcot et al. (1989) states that confirmation bias "was manifested principally as a weighting and not as a distortion bias." Further, the Cheikes article indicates that "ACH had no impact on the weight effect," having tested both elicitations of the bias (in both evidence selection and evidence weighting). 
Emily Pronin (2007), the leading authority on Bias Blind Spot, presents a similar conclusion: "Participants not only weighted introspective information more in the case of self than others, but they concurrently weighted behavioral information more in the case of others than self."
Robert Jervis, professor of International Affairs at Columbia University, speaks about evidence-weighting issues in the context of the Fundamental Attribution Error in his 1989 work Strategic Intelligence and Effective Policy.
What if the impact of bias in analysis is less about deciding which pieces of evidence to use and more about deciding how much influence to allocate towards each specific piece?  This would mean that to mitigate the effects of cognitive bias and to improve forecasting accuracy, training programs should focus on teaching analysts how to weight and assess critical pieces of evidence.

With that question in mind, I designed a simple experiment with four distinctly testable groups to assess the effects of evidence weighting on a) cognitive bias and b) forecasting accuracy. 

Each of the four groups were required to spend approximately one hour conducting research on the then-upcoming Honduran presidential election to determine a) who was most likely to win and b) how likely they were to win (in the form of a numerical probability estimate, e.g. "X is 60 percent likely to win"). Each group, however, used varying degrees of Analysis of Competing Hypotheses (ACH), allowing me to manipulate how much or how little the participants could weight the evidence. A description of each of the four groups is below:
  • Control group (Cont, N=28). The control group was not permitted to use ACH at all. They had one hour to conduct research independently with no decision support tools. 
  • ACH no weighting (ACH-NW, N=30). This group used ACH.   Participants used the PARC 2.0.5 ACH software without the ability to use II (highly inconsistent) or CC (highly consistent) functions. Nor were they allowed to use the credibility or the relevance functions.
  • ACH with weighting (ACH-W, N=30). This group used ACH as they had been instructed, including II, CC and relevance, but not credibility.
  • ACH with training (ACH-T, N=27). This was the focus group for the experiment. Participants in this group, which used ACH with full functionality (excluding credibility), first underwent a 20-minute instructional session on evidence weighting and source reliability employing the Dax Norman Source Evaluation Scale and other instructional material. In other words, these participants were instructed how to weight evidence properly. 
While the election prediction served as the metric for assessing forecasting accuracy (the experiment was conducted two weeks before the election), five separate instruments were used in the form of a post-test in order to elicit bias, three of which corresponded with confirmation bias, one which addressed the framing effect and then, finally, representativeness. 

The results were intriguing:

The group with the most accurate forecasts (79 percent) was the control group, or the group that did not use ACH at all (See Figure 1). The next most accurate group (65 percent) was the ACH-T group, or the ACH "with training." 


Figure 1. The Effects of Evidence Weighting Across Four Groups on Forecasting Accuracy and Cognitive Bias.
Note: The percentage for each bias represents the percentage of unbiased responses obtained in that group.
Due to the small sample sizes, these differences did not turn out to be statistically significant which, in turn, suggests the first major point:  That training in cognitive bias mitigation and some structured analysis techniques might not be as useful as originally thought.  

If this were the first time these kinds of results had been found, it might be possible to chalk it up to some sort of sampling error.  But Drew Brasfield found much the same thing when he ran a similar experiment back in 2009  (The relevant charts and texts are on pages 38-39).  In Brasfield's case, participants were statistically significantly less biased when they used ACH but forecasting accuracy remained statistically significantly the same (though, in this experiment, the ACH group technically outperformed the Control).

These results also suggest that more accurate estimates came from analysts who either a) intuitively weighted evidence without the help of a decision tool or b) were instructed how to use the decision tool with special focus on diagnosticity and evidence weighting. This could mean that analysts, when given the opportunity to weight evidence without knowing how much weighting and diagnosticity impacts results, weight incorrectly out of perceived obligation to do so or misunderstanding.

Finally, the next lowest forecasting accuracy was obtained by the group ACH-NW (53 percent) in which the analysts were not allowed to weight evidence at all (no IIs or CCs). The lowest accuracy (only 45 percent) was obtained by the group that was permitted to weight evidence with the ACH decision tool but were not instructed how to do so nor were they informed how this weighting might influence the final inconsistency scores of their hypotheses.  This final difference was statistically significant from the control suggesting that a failure to train how to weight evidence appropriately actually generates lower forecasting accuracy.

If that weren't enough, let's take one more interesting look at the data...

In terms of analytic accuracy, the hierarchy is as follows (from most to least accurate): Control, ACH-T, ACH-NW, ACH-W.

Now, in terms of most biased, the hierarchy looks something like this (from least to most biased):
  • Framing: ACH-T, ACH-W, ACH-NW, Control
  • Confirmation: ACH-W, ACH-T, Control, ACH-NW
  • Representativeness: ACH-W, ACH-NW, Control, ACH-T
What this shows is an (albeit imperfect) inverse to analytic accuracy. In other words, the more accurate groups were also more biased, and while ACH generally helped mitigate bias, it did not improve forecasting accuracy (in fact, it may have done the opposite). If this experiment achieved its goal and effectively measured evidence weighting as an underlying mechanism of forecasting accuracy and cognitive bias, it supports the claim made by Cheikes et al. above: "ACH had no impact on the weight effect" (again, talking about confirmation bias) and, as mentioned, recreates the results found by Brasfield. 

While the evidence weighting hypothesis is obviously in need of further investigation, this preliminary experiment provided initial results with some intriguing implications, the most impactful of which is that, while the use of ACH reduces the effects of cognitive bias, it may not improve forecasting accuracy. A less biased forecast is not necessarily a more accurate forecast. 



***
As a side note, I wanted to include this self-reported data which shows the components that the 115 analysts in this experiment indicated were most influential in their final analytic estimates in general. Note that they indicate that source reliability and availability of information seem to be the top two (See Figure 2).

Figure 2. Self-Reported Survey Data of 115 Analysts Indicating Factors That Most Influence Their Analytic Process
Scale = 1 - 4

REFERENCES

Cheikes, B. A., Brown, M. J., Lehner, P. E., & Adelman, L. MITRE, Center for integrated intelligence systems. (2004). Confirmation bias in complex analyses (51MSR114-A4). Bedford, MA.

Jervis, R. (1989). Strategic intelligence and effective policy. In Frank Cass (Ed.), Intelligence and security perspectives for the 1990s. London, UK.

Nickerson, R. S. (1998). Confirmation bias: An ubiquitous phenomenon in many guises. Review of general psychology, 2(2), 175-220.

Pronin, E., Gilovich, & Ross, L. (2002). The bias blind spot: Perceptions of bias in the self versus others. Personality and social psychology bulletin, 28, 369-381.

Pronin, E., & Kugler, M. B. (2007). Valuing thoughts, ignoring behavior: The introspection illusion as a source of the bias blind spot. Journal of experimental social psychology, 43, 565-578.

Tolcott, M. A., Marvin, F. F., & Lehner, P. E. (1989). Expert decisionmaking in evolving situations. IEEE transactions on systems, man, and cybernetics, 19(3), 606-615.

Friday, August 13, 2010

Does Analysis Of Competing Hypotheses Really Work? (Thesis Months)

The recent announcement that collaborative software based on Richards Heuer's famous methodology, Analysis of Competing Hypotheses, would soon be open-sourced was met with much joy in most quarters but some skepticism in others.

The basis for the skepticism seems to be the lack of hard evidence that ACH actually improves forecasting accuracy.  While this was not the only (and may not have been the most important) reason why Heuer created ACH, it is certainly a question that bears asking.

No matter how good a methodology is at organizing information or creating an analytic audit trail or easing the production burden, etc., the most important element of any intelligence methodology would seem to be its ability to increase the accuracy of the forecasts generated by the method (over what is achievable through raw intuition). 

With a documented increase in forecasting accuracy, analysts should be willing to put up with almost any tedium associated with the method.  A methodology that actually decreases forecasting accuracy, on the other hand, is almost certainly not worth considering, much less implementing.  Methods which match raw intuition in forecasting accuracy really have to demonstrate that the ancillary benefits derived from the method are worth the costs associated with achieving them.

It is with this in mind that Drew Brasfield set out to test ACH in his thesis work while here at Mercyhurst.  His research into ACH and the results of his experiments are captured in his thesis, Forecasting Accuracy And Cognitive Bias In The Analysis Of Competing Hypotheses (full text below or you can download a copy here).

To test ACH, Drew used 70 students divided between a control and an experimental group who were all familiar with ACH.  The groups were asked to research and estimate the results of the 2008 Washington State gubernatorial election between Democrat Christine Gregoire and Republican Dino Rossi (Gregoire won the election by about 6 percentage points).  The students were given a week in September 2008 to independently work on their estimate of who would win the election in November.

The results were in favor of ACH in terms of both forecasting accuracy and bias.  In Drew's words, "The findings of the experiment suggest ACH can improve estimative accuracy, is highly effective at mitigating some cognitive phenomena such as confirmation bias, and is almost certain to encourage analysts to use more information and apply it more appropriately."

The results of the experiment are displayed in the graphs below:
Statistical purists will argue that the results did not meet the traditional 95% confidence interval test suggesting that the accuracy difference may be due to chance. True enough. What is clear, though, is that ACH doesn't hurt forecasting accuracy and, when combined with the other results from the experiment (see below) strongly suggests that Drew's characterization of ACH is correct.

Becasue Drew captured the political affiliation of his test subjects before he conducted his experiment he was able to sort those subjects more or less evenly into the control and experimental groups.  Here again, ACH comes away looking pretty good:
The chart may be a bit confusing at first but the bottomline is that Republicans were far more likely to accurately forecast the eventual victory of the Democratic candidate if they used ACH.  Here again the statistics suggest that chance might play a larger role than normal (an effect exacerbated by the even smaller sample sizes for this test).  At the least, however, these results are consistent with the first set of results and, again, do nothing to suggest that ACH does not work.

Drew's final test is the one that helps clarify any fuzziness in the results so far.  Here he was looking for evidence of confirmation bias -- that is, analysts searching for facts that tend to confirm their hypotheses instead of looking at all facts objectively.  He was able to find statistically significant amounts of such bias in the control group and almost none in the experimental group:
It is difficult for me to imagine a method which worked so well at removing biases that would also not improve forecasting accuracy. In short, based on the results of this experiment, concluding that ACH doesn't improve forecasting accuracy (due to the statistical fuzziness) would also require one to conclude that biases don't matter when it comes to forecasting accuracy. This is an arguable hypothesis, I suppose, but not where I would put my money...

The most interesting part of the thesis, in my opinion, though, is the conclusion.  Here Drew makes the case that the statistical fuzziness was a result of the kind of problem tested, not the methodology.  He suggests that "ACH may be less effective for an analytical problem where the objective probabilities of each hypothesis are nearly equal."

In short, when the objective probability of an event approaches 50%, ACH may no longer have the resolution necessary to generate an accurate forecast.  Likewise, as objective reality approaches either 0% or 100%, ACH becomes increasingly less necessary as the correct estimative conclusion is more or less obvious to the "naked eye". Close elections, like the one in Washington State in 2008 may, therefore, be beyond the resolving power of ACH.

Like much good science, Drew's thesis has generated a new testable hypothesis (one we are, in fact, in the process of testing!).  It is definitely worth the time it takes to read.

Forecasting Accuracy and Cognitive Bias in the Analysis of Competing Hypotheses

Friday, December 19, 2008

Top 5 Intelligence Analysis Methods: Analysis Of Competing Hypotheses (#1)

(Note: I apologize for how long it has taken me to get here. Conferences, classes and life in general conspired to get in the way this week. For the patient, here is the last in this series of posts...)

Part 1: Introduction
Part 2: What Makes A Good Method?
Part 3: Bayesian Analysis (#5)
Part 4: Intelligence Preparation Of The Battlefield/Environment (#4)
Part 5: Social Network Analysis (#3)
Part 6: Multi-Criteria Decision Making Matrices/Multi-Criteria Intelligence Matrices (#2)

Analysis Of Competing Hypotheses (ACH) is probably the best-known intelligence analysis method today. Invented by Richards Heuer over 30 years ago and made famous in his intelligence classic, the Psychology of Intelligence Analysis, ACH is widely taught and conceptually easy even for entry-level analysts.

In addition, it was specifically designed to work in all kinds of situations with any kind and quality of data. What is less clear is that the method produces unequivocally better estimative results. While the method is rooted fundamentally in the scientific method, studies testing the value of the method as a way to improve forecasting have been few and the results have been mixed (In no study have the results been worse than without the method but some studies have shown that the method only helps certain subsets of analysts. For a good list of these studies see the Notes List at the end of the ACH article on Wikipedia).

I am not sure why this is so. My own impression is that a well done ACH provides a better estimate in less time and with more nuance than virtually any other method available.

We teach ACH here in our freshman classes. I see many, many students struggle not with the basic concepts of ACH but with the details. I see countless examples each year of student projects where they have improperly executed the method (in much the same way a student gets their first attempts at a calculus or chemistry problem incorrect).

In most cases, it is fairly easy to correct the mistakes and the students rarely have a problem seeing what they did wrong or in making the appropriate adjustments. It is less clear to me that, at this early stage in their education, they are able to transfer this knowledge from one type of problem to another, however. We try to reinforce all our methods in upper level classes but the opportunities for reinforcement in the real world are slim (we rarely find, for example, that students are required to use structured methods in their internships).

My own instincts tell me that ACH (and many of the experiments involving it -- including our own) is a powerful method but won't get a fair test until such a test is done with analysts who have worked with the method on multiple problems and in multiple circumstances. To be honest, I suspect that this is true with all of the methods I have discussed in this series. Deliberate practice seems to be a key component of expertise in multiple other fields and I imagine this is true when it comes to intelligence analysis methods as well.

Improving the quality of the final estimate is only one (albeit an important) way that a method should contribute to a quality intelligence product, however. ACH brings much more to the table in my estimation and it does this immediately, in even the earliest projects.

ACH can help the analyst at every stage of the problem, including modelling, collection and collection planning, and preparing a document for dissemination. It is a wholly transparent method and can very easily be used collaboratively. Its transparency is also crucial in helping instructors or managers identify problems in the analysis of the data. The transparency is also of enormous benefit in understanding and improving the analytic process after the fact as well. It integrates extremely well with various data resources and is very suitable for automation. We find that it is actually faster to use, particularly in a group setting, than most other methods (including intuitive analysis).

The way ahead is a little different here than with the other methods. We think we have a pretty good handle on how to teach ACH. The key, in my estimation, is to create opportunties to reinforce that teaching in and outside the confines of the classroom.