Friday, April 10, 2009

Boxes That Misunderstand Themselves

The explanations that some putative experts give for their decisions may not be accurate; i.e., those given reasons may not explain why those putative experts decide as they do (or say what they say). If that is the case, understanding these experts' stated reasons would not allow us to predict the decisions (or inferences) of these supposed experts. However, it does not follow that such putative experts do not follow some rules or principles (that are unknown to them). In short, sometimes we will distrust the explanations that some experts give but we may yet believe that these supposed experts will sometimes be useful barometers (for reasons they themselves do not accurately understand). But to figure out just when these nincompoopish experts will be useful barometers and when they won't, we need to understand what makes them tick -- what really makes them tick. Otherwise we may foolishly say, "Well, these people correctly predicted the last recession. Even if they can't explain what leads them to make the predictions they do, it's a good bet they'll correctly predict the next recession." Well, maybe, and maybe not.

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

B.F. Skinner's Rats & Pigeons in B.F. Skinner's Mazes

B.F. Skinner (the famous behaviorist) wasn't interested in the internal mechanisms of the rats and pigeons in his mazes. B.F. Skinner was interested only (he said, as I recall) in the responses of his animals to positive and negative inputs (rewards).

But B.F. Skinner's animals would not have responded the way they did in the past if someone had snipped the chains of neurons and axons (or whatnot) that transmitted sensory signals from the animals' environments to the innards of the animals that B.F. Skinner (said he) didn't care about.

So what does this have to do with proficiency testing of experts?

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Predicting (Inferring) a Box's Behavior (Outputs, Reports)

Assume:
Input: an observation (a signal, a possibly-sensed event).

Output: a statement.

Intermediary: a box.

Question: To predict a box's outputs given specified inputs, do you have to be able to see (or infer) the innards (or workings) of the box or is it sufficient to be able to observe the box's outputs in the past given specified inputs?

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

There Is No Law against Blue Skies, Is There?

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Evidence of Things to Come

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Spring: The Slow-Thinking Season

Legal business, law schools, and analytical thinking all slow during the "spring break," while everyone (here in the Northeast, in any event) impatiently awaits the arrival of warm & sunny weather. (But migratory ducks are already here and are almost gone.)

Well, in recompense perhaps I'll post something tonight or tomorrow about measuring the proficiency of putative experts. Otherwise I'll post a digital image or two.

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Sunday, March 29, 2009

New Jersey Thinks Again about the Polygraph Test

The Supreme Court of New Jersey recently declined to completely outlaw the admission of polygraph evidence. However, the NJ Supreme Court retained its rule that in the absence of a stipulation, polygraph evidence is inadmissible. In addition, it held that a stipulation without the advice of counsel is ineffective. Finally, it said that the "next time" a party seeks to introduce polygraph evidence pursuant to a stipulation, the trial court must hold a hearing to determine the reliability of polygraph evidence. See State v. A.O., --- N.J. ----, --- A.2d ----, 2009 WL 529149 (N.J.,March 4, 2009). In reaching these conclusions New Jersey's Supreme Court said (footnotes omitted) the following things:

IV.

We next consider the enforceability of the stipulation in light of the law regarding polygraph evidence and the facts of this case.

A.

As a general rule, polygraph results are not admissible in evidence in New Jersey. State v. Domicz, 188 N.J. 285, 312-13, 907 A.2d 395 (2006); McDavitt, supra, 62 N.J. at 44, 297 A.2d 849; State v. Driver, 38 N.J. 255, 261, 183 A.2d 655 (1962). In 1972, this Court held in McDavitt that "to date ... lie detector testing has not yet attained scientific acceptance as a reliable and accurate means of ascertaining truth or deception." 62 N.J. at 44, 297 A.2d 849. We reaffirmed that view recently and noted that "[i]n the more than thirty years since McDavitt, serious questions about the reliability of polygraph evidence remain." Domicz, supra, 188 N.J. at 313, 907 A.2d 395.

There remains a "lack of scientific consensus concerning the reliability of polygraph evidence, which in turn is reflected in the disagreement among state and federal courts concerning the admissibility of such evidence." Id. at 312, 907 A.2d 395 (citing United States v. Scheffer, 523 U.S. 303, 309-12, 118 S.Ct. 1261, 1265-66, 140 L.Ed.2d 413, 419-21 (1998) (reviewing scientific studies showing that accuracy of polygraph tests ranges from 50 to more than 90 percent)). Some studies suggest that the accuracy rate is "little better than could be obtained by the toss of a coin." Scheffer, supra, 523 U.S. at 310, 118 S.Ct. at 1265, 140 L.Ed.2d at 419 (citing Iacono & Lykken, The Scientific Status of Research on Polygraph Techniques: The Case Against Polygraph Tests, in 1 Modern Scientific Evidence § 14-5.3).

Nonetheless, to many citizens who serve on juries, polygraph evidence-- presented by experts and arrayed in scientific language--has an aura of infallibility. That impression "can lead jurors to abandon their duty to assess credibility and guilt" and rely instead on the examiner's expert opinion. Scheffer, supra, 523 U.S. at 314, 118 S.Ct. at 1267, 140 L.Ed.2d at 422. As a result, "the vast majority of states either ban polygraph evidence altogether or do not admit such evidence absent a stipulation between the State and defendant." Domicz, supra, 188 N.J. at 312-13, 907 A.2d 395.

Twenty eight states bar the admission of polygraph evidence outright. ...; see also People v. Angelo, 88 N.Y.2d 217, 644 N.Y.S.2d 460, 666 N.E.2d 1333, 1335 (1996) (polygraph evidence properly excluded where there continues to be no showing that such evidence is generally accepted as reliable by scientific community).

Virtually all the other states to consider the issue--eighteen in total--limit the admission of polygraph evidence to cases where both parties stipulate to its use. ...

Only New Mexico allows the admission of polygraph exam results without stipulation. Lee v. Martinez, 136 N.M. 166, 96 P.3d 291, 306-07 (2004).

Underscoring the widespread skepticism about the polygraph's reliability, four states--Massachusetts, Wisconsin, North Carolina, and Oklahoma--have experimented with allowing the admission of polygraph evidence for a number of years, only to reject the practice and reinstate the traditional rule of inadmissibility. ...

Our view remains unchanged. This Court has not sanctioned and does not now entertain the admission of polygraph results. Nor does our holding in McDavitt offer support for the admission of the stipulated polygraph results in this case. That holding addressed very different facts, and we once again decline to "'widen the small aperture of ... McDavitt.' " State v. Baskerville, 73 N.J. 230, 236, 374 A.2d 441 (1977) (quoting State v. Cole, 131 N.J.Super. 470, 471, 330 A.2d 594 (App.Div.1974)); see also Domicz, supra, 188 N.J. at 313, 907 A.2d 395 ("[W]e are not prepared to extend McDavitt to unstipulated polygraph examinations, even in a suppression hearing presided over by a judge.").

McDavitt created a very narrow exception to the rule barring polygraph evidence. In that case, the defendant's conduct before the jury provoked the defensive use of polygraph evidence. During his criminal trial, the defendant testified that, after his arrest, he had offered to take a polygraph test to prove his innocence. McDavitt, supra, 62 N.J. at 41, 297 A.2d 849. The prosecutor objected and was mistakenly overruled. *163 Id. at 41, 43, 297 A.2d 849. With the door thus opened, the prosecutor asked on cross-examination if the defendant would be willing to take a polygraph that day. He was. Id. at 41, 297 A.2d 849. After further discussion outside of the jury's presence, the trial court granted a recess to allow the defendant time to confer with his lawyer. Id. at 42, 297 A.2d 849. Afterward, with the court's approval, the parties stipulated as follows: if the defendant passed the test, the State would not oppose a motion for acquittal; if he failed, the test results would be presented to the jury. Id. at 41-42, 297 A.2d 849.

Those unusual facts gave rise to the exception the Court framed: polygraph results may be admitted in evidence on agreement of the parties if their stipulation is "clear, unequivocal and complete, freely entered into with full knowledge of the right to refuse the test and the consequences involved in taking it." Id. at 46, 297 A.2d 849. In addition, the examiner must be qualified and the test administered in accordance with established techniques. Ibid.

McDavitt neither discussed nor sanctioned a polygraph stipulation agreed to by a suspect alone. McDavitt, therefore, does not offer support for the stipulation used in this case.

B.

We are troubled by more than the prosecution's misplaced reliance on McDavitt and have concerns about certain matters defendant was asked to stipulate to on his own.

First, we question defendant's ability to stipulate to the expert's qualifications. Defendant acknowledges in the stipulation that the polygrapher was an "expert in all phases of both administering polygraph examinations and in the analysis of polygraph chart recordings." How can a suspect, unschooled in the complexities of polygraphy or the credentials needed to administer a valid examination, stipulate to that statement? What factual basis does a suspect possess to form a view of the examiner's expertise? Nothing in the record allays this concern. As the Appellate Division noted, "[i]f this were a consumer contract, we might deem it unconscionable." A.O., supra, 397 N.J.Super. at 23, 935 A.2d 1202.

Second, the stipulation waives all challenges to the admissibility of the polygraph expert's testimony. Although defendant may cross-examine the expert about his or her qualifications, the manner in which the examination was conducted, the expert's opinion, and the possibility of error, the stipulation nonetheless provides for the automatic right of the expert to testify. In other words, even if defense counsel can undermine basic foundational elements of the expert's testimony and establish at trial that the polygrapher was wholly unqualified, the opinion voiced was not well-grounded, or that the possibility of error was great, the stipulation authorizes the expert to present his or her findings to the jury. That practice offends the core purpose of our evidentiary rules. See N.J.R.E. 403, 702.

Third, the stipulation limits defendant's ability to attack the polygraph evidence. While he may cross-examine the State's expert, defendant cannot call another witness on the subject. In other words, another expert, no matter how well qualified, cannot offer a contrary opinion about the test results. From accident reconstruction to blood-sample tests, it is common practice for a party to try to rebut the other side's expert testimony with an expert of its own. To be sure, we have strong reservations about allowing dueling experts to testify about polygraph results because of doubts about the polygraph's reliability in general. See Domicz, supra, 188 N.J. at 314, 907 A.2d 395. But in our adversary system of justice, that legal issue is best addressed by lawyers, not suspects.

Fourth, the stipulation collapses questions about a suspect's voluntary consent with the legal issue of admissibility. In evaluating a waiver of rights, the focus at first is on whether a defendant knowingly and voluntarily entered into the waiver agreement. Next, the focus shifts to whether the results of that waiver may be admitted in evidence. For example, a defendant can knowingly consent to a search, but in doing so does not agree to the admissibility of everything found during the search. The State must still establish that the evidence taken is admissible in accordance with substantive and evidentiary rules. A seized document that would otherwise be inadmissible--whether because the material was irrelevant, prejudicial, privileged, or hearsay--is not cured of its inadmissibility simply because a citizen agreed to its seizure. See, e.g., N.J.R.E. 401, 403, 504, 702. Likewise, defendants may waive their Miranda rights, but they do not stipulate to the admission of all statements that follow. An irrelevant or highly prejudicial comment would still be subject to evidentiary rules that might bar such statements. The same is true for a polygraph exam. A defendant can voluntarily agree to take the test, but its admissibility is a distinctly separate question.

Once properly advised of his rights, defendant could agree to submit to a polygraph. But the ancillary decisions made beyond that choice bear on trial strategy. Defendants typically rely on counsel to object to otherwise inadmissible evidence, attack a witness's expertise, and decide the most effective way to challenge evidence before a jury. See Rules of Professional Conduct 1.2 (allocating authority between lawyer and client). The stipulation here, though, operated to eliminate counsel's role by relying on a suspect's consent.

To avoid that course, a number of other states allow polygraph results by stipulation only upon the approval of defendant's counsel. ... Such an approach is consistent with the holding in McDavitt but was not followed here.

Our "overarching constitutional responsibility [is] to guarantee the proper administration of justice." State v. Williams, 93 N.J. 39, 62, 459 A.2d 641 (1983). "When we perceive ... that more might be done to advance the reliability of our criminal justice system, our supervisory authority over the criminal courts enables us constitutionally to act." State v. Romero, 191 N.J. 59, 74-75, 922 A.2d 693 (2006) (citing N.J. Const. art. VI, § 2, ¶ 3; State v. Delgado, 188 N.J. 48, 62, 902 A.2d 888 (2006)). We do so now to ensure greater fairness at trial and reliability of jury verdicts.

Relying on our supervisory authority, we bar the introduction of polygraph evidence based on stipulations entered into without counsel. We therefore affirm the Appellate Division's decision to reverse defendant's conviction. The conviction rested on the testimony of a young witness who recanted and then withdrew her recantation. No physical or medical evidence corroborated her testimony. To strengthen its case, the State introduced and highlighted the polygraph evidence discussed above and presented it as "100 percent accurate." We agree with the Appellate Division that "the polygraph evidence may well have made the difference between conviction and acquittal in this case." A.O., supra, 397 N.J.Super. at 33-34, 935 A.2d 1202. As a result, admission of the evidence was clearly capable of producing an unjust result, see Rule 2:10-2, and warrants reversal and a new trial.

C.

Judge Weissbard's concurring opinion [in the opinion in this case of New Jersey's intermediate appellate court] encourages us to take one more step: to reverse McDavitt and ban polygraph evidence altogether. He reminds us that the core concern of our evidence rules "is to provide the fact-finder with only reliable and probative evidence." A.O., supra, 397 N.J.Super. at 30, 935 A.2d 1202 (Weissbard, J.A.D., concurring) (citing 1 Wigmore on Evidence § 7a (Tillers rev.1983)); see also Scheffer, supra, 523 U.S. at 309, 118 S.Ct. at 1264, 140 L.Ed.2d at 419. As was true in Domicz, however, we do not have an adequate record to make ultimate findings about the reliability of polygraph evidence at this time. See Domicz, supra, 188 N.J. at 312-13, 907 A.2d 395. Nonetheless, we harbor a number of concerns about McDavitt in light of developments since 1972.

McDavitt recognized that "lie detector testing has not yet attained scientific acceptance as a reliable and accurate means of ascertaining truth or deception," but concluded that the "art of polygraph testing had developed to a point that its results were probative enough to warrant admissibility upon stipulation." 62 N.J. at 44, 297 A.2d 849. As support for that finding, McDavitt cited to two criminal trial courts that had conducted extensive hearings on the reliability of polygraph tests and found that the results were "now generally accepted by authorities in the field and ... capable of producing highly probative evidence in a court of law when properly used by competent, experienced examiners." Id. at 45, 297 A.2d 849. The Court cited specifically to United States v. Ridling, 350 F.Supp. 90 (E.D.Mich.1972), and United States v. Zeiger, 350 F.Supp. 685 (D.D.C.1972). Zeiger, however, was reversed summarily. See United States v. Zeiger, 475 F.2d 1280 (D.C.Cir.1972). And Ridling was later criticized by its own and one other circuit court for its treatment of polygraph evidence. United States v. Alexander, 526 F.2d, 161, 166 (8th Cir.1975); United States v. Frogge, 476 F.2d 969, 970 (5th Cir.1973).

Furthermore, as discussed above, after 1972 four states allowed the admission of polygraph evidence for a number of years but reversed course because of questions about reliability among other reasons. See Commonwealth v. Mendes, supra, 547 N.E.2d at 41; Dean, supra, 307 N.W.2d at 653; Grier, supra, 300 S.E.2d at 359-60; Fulton, supra, 541 P.2d at 872; see also Porter, supra, 698 A.2d at 775-76.

Recent social science studies cast doubt on the reliability of polygraph evidence as well. See Scheffer, supra, 523 U.S. at 309-10, 118 S.Ct. at 1265, 140 L.Ed.2d at 419-20 (reviewing social science evidence); Porter, supra, 698 A.2d at 759-68 (same); NRC Study, supra, at 323-53 (2003) (reviewing 194 separate studies of polygraph testing).

Those studies explain that polygraphy relies on two assumptions: (1) that deception triggers certain emotional states; and (2) that those emotional states produce specific, measurable physiological changes in the body. Porter, supra, 698 A.2d at 759. As certain empirical evidence has shown, however, there is substantial variation in how individuals respond physiologically when they are lying or telling the truth, and the responses that humans produce in such situations are not specific to either deception or truth-telling. Id. at 760 (citations omitted); NRC Study, supra, at 212-13. The inherent ambiguities in such responses, which arise from individual variations in the subject's cardiovascular, electrodermal and respiratory activity, often make it difficult for a test administrator to determine if the examinee is lying, nervous, tired, or simply trying to game the system. NRC Study, supra, at 4, 13-17, 216, 286-90.

As the Supreme Court [of the United States] observed, "there is simply no way to know in a particular case whether a polygraph examiner's conclusion is accurate, because certain doubts and uncertainties plague even the best polygraph exams." Scheffer, supra, 523 U.S. at 312, 118 S.Ct. at 1266, 140 L.Ed.2d at 421. Even more troubling, "to the extent that the polygraph errs, studies have repeatedly shown that the polygraph is more likely to find innocent people guilty than vice versa." Id. at 333, 118 S.Ct. at 1276, 140 L.Ed.2d at 433-34 (Stevens, J., dissenting).

Compounding these questions about reliability is the fact that many lay people tend to view polygraph evidence as bordering on infallible. Id. at 314, 118 S.Ct. at 1267, 140 L.Ed.2d at 422 (majority opinion); A.O., supra, 397 N.J.Super. at 33, 935 A.2d 1202 (citations omitted). Thus, potentially unreliable polygraph evidence may receive undue weight and distract jurors from judging the credibility of witnesses directly.

Such concerns raise questions about the continuing wisdom of McDavitt. Because we lack a factual record, we cannot fully address those issues today. However, a proper record will have to be developed in the trial court the next time a party seeks to introduce stipulated polygraph evidence, agreed to by both sides. That evidence should be introduced only if the parties can first establish its reliability at an N.J.R.E. 104 hearing.

END OF OPINION

Notes by Peter Tillers:

1. The Supreme Court of New Jersey is unwilling to extend the principle of party autonomy to allow the outcome of a trial to be determined or affected by the equivalent of a coin toss -- even if the parties so stipulate. Hurrah for the Supreme Court of New Jersey!

2. When a trial court considers the question of the reliability of polygraph testing in the next New Jersey case in which a party seeks to introduce polygraph evidence pursuant to a stipulation, it is unlikely that the trial court will find that polygraph evidence is "reliable" even if the test is administered under ideal conditions.

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Friday, March 27, 2009

Postscript to the Story of the DNA Pixie Dust in the Phantom of Heilbronn Case

A representative of the state prosecutor's office in Heilbronn officially confirmed that the source of the DNA found at about 40 crime scenes in Europe was not the hypothesized persistent peripatetic perpetrator called the "Phantom of Heilbronn" but, rather, a mercifully-unnamed female employee at an unnamed packing company somewhere in Bavaria. (The female employee presumably did not bother to wear gloves when packing the cotton swabs.) See "'Phantom-Mörderin' ist ein Phantom [The Phantom Murderer is a Phantom]," Spiegel Online (March 27, 2009).

The "Gold Standard" Bites the Dust! (gratuitous editorial comment by PT)

So: Did someone win the 300,000 Euro award for providing definitive clues to the (non-)identity of the Phantom of Heilbronn? If not, why not?


the dynamic evidence page




Thursday, March 26, 2009

DNA as Pixie Dust

For many months police in Germany and other parts of Europe have been searching for the "phantom of Heilbronn." See my earlier post, A Malicious Fairy Princess Spreading Spreading Pixie Dust Laced with DNA over European Crime Scenes? (May 14, 2008). This "woman without a face," DNA tests showed, had committed six different murders and also other crimes in various towns in Germany and Europe. A reward of 300,000 Euros was offered for clues to the identity of this female miscreant.

Oops. Yes, the crimes were committed. But the thesis that one evil woman committed all of them was, it turns out, very, very probably false. See "Eine sehr peinliche Geschichte," Spiegel Online (March 26, 2009).

What led to this snafu?

DNA.

More precisely, unsanitary packing procedures during the manufacturing process, perhaps.

The Spiegel story (in very rough translation) tells us:
It now appears that the chances of finding the trail of the mysterious suspect are slimmer than ever before -- because she apparently doesn't exist. According state's attorney Heilbronn, the Baden-Wuerttenberg Office of Criminal Affairs is now investigating whether the cotton swabs ["Q-tips"] that investigators used to collect DNA samples [at crime scenes] had already been contaminated with DNA and had thus led investigators on a false trail.

According to a report in Stern.de, the matter involves a packager [a packer, an employee] who worked for the manufacturer of the swabs that were used [at the crime scenes]. According to this report, the swabs were, to be sure, sterilized. But, according to Christian Rueff of the University of Zurich, such sterilization of the swabs does not affect contamination of the swabs with DNA by cells of the human body. [PT: The employee, I presume, held the cotton swabs in her hands when putting them into boxes or other containers.] The manufacturer delivered cotton swabs to various places in Germany and and also in France and Austria.
The article then proceeds to describe what led some observers to believe that the phantom woman might not in fact exist. One general problem was the large number of crimes -- 40 -- that the "phantom" supposedly committed in, supposedly, widely dispersed locations in Europe. Another hint that something might be amiss was that the supposed perpetrator supposedly began committing the crimes in 1993. But the immediate reason for doubt that a single woman committed the 40 or so crimes was the discovery that swabs of documents owned by a person who died in a fire had DNA that matched the DNA of the "phantom" -- which, an official proclaimed, "really could not be."

An official at the state prosecutor's office stated it was even possible that the cotton in the cotton swabs used in the police investigations was contaminated with DNA when it was plucked from the plant.


&&&

My thanks to Lothar Philipps for alerting me to this story.


&&&

Clarification: The corpse found in the fire apparently was never thought to be the corpse of the suspected culprit. The testing of the documents the dead person had owned aroused suspicion because the second time the documents were tested no matching DNA was found. See "'DNA bungle' haunts German police," BBC News (March 26, 2009).


the dynamic evidence page












Wednesday, March 25, 2009

Folk physics, common sense, inference, and intelligence

Perhaps epistemology (a general theory of human knowledge) must ride on the back of ontology (a theory of what is and how things-that-are work).

Perhaps metaphysics = sophisticated folk physics.

Folk physics should make use of the insights of physics -- and much else (e.g., neuroscience).

But folk physics -- metaphysics -- should not allow itself to be displaced by, e.g., physics or neuroscience.

Do not think that matters such as common sense and folk physics are unintelligent. They harbor much intelligence. If a special science can explain them, it will prove that it too is very intelligent. (But no special science can yet explain common sense, or "ordinary" intelligence.)

Of course, physics, neuroscience, etc., have much intelligence that ordinary intelligence lacks. But this fact does not render folk physics, common sense, etc., unintelligent.

&&&

the dynamic evidence page

coming soon: the law of evidence on Spindle Law

Hájek: Whether You Like It or Know It or Not, You Have a Reference Class Problem

Alan Hájek is an extraordinarily perceptive commentator on (the) reference class problem(s). See "The Reference Class Problem is Your Problem Too", Synthese 156: 185-215. 2007. He maintains that conditional probability should be a "primitive" for axiomatization of probability theory and that taking this view dissolves the "metaphysical" form of (the) reference class problem(s). However, he candidly admits this approach does not dissolve the epistemological version of the problem(s) [footnotes omitted]:
Now, the bad news. Giving primacy to conditional probabilities does not so much rid us the epistemological reference class problem as give us another way of stating it. Which of the many conditional probabilities should guide us, should underpin our inductive reasonings and decisions? Our friend John Smith is still pondering his prospects of living at least eleven more years as he contemplates buying life insurance. It will not help him much to tell him of the many conditional probabilities that apply to him, each relativized to a different reference class: “conditional on your being an Englishman, your probability of living to 60 is x; conditional on your being consumptive, it is y; …”. (By analogy, when John Smith is pondering how far away is London, it will not help him much to tell him of the many distances that there are, each relative to a different reference frame.) If probability is to serve as a guide to life, it should in principle be possible to designate one of these conditional probabilities as the right one. To be sure, we could single out one conditional probability among them, and insist that that is the one that should guide him. But that is tantamount to singling out one reference class of the many to which he belongs, and claiming that we have solved the original reference class problem. Life, unfortunately, is not that easy—and neither is our guide to life.

Still, it’s better to have one problem than two. I will leave it to others to judge the extent to which I have succeeded in ridding us of the metaphysical reference class problem. But I am aware that I have not solved the epistemological problem. I invite you to join me in the search for a solution for the interpretations of probability that have a genuine claim to being guides to life. After all, whichever interpretation you favor, the epistemological version of the reference class problem is your problem too.

  • One man's metaphysics is another man's physics?
    Folk physics?
    Folk physics has its uses -- and in an important sense it may be even "true."
  • P.S. I suspect that a successful "solution" to the epistemological version of the reference class problem(s) requires a bit of metaphysics (i.e., some basic assumptions about [wo]man and the world [s]he inhabits).
    N.B. By putting solution in quotation marks I do not mean to assert that a satisfactory solution of some kind is impossible. (Of course, there are solutions and there are solutions: one solution will not necessarily solve -- thank goodness! -- every possible problem about the choice and use of reference classes.)

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Tuesday, March 24, 2009

    A Practical Solution to the Reference Class Problem?

    Professor Edward K. Cheng argues for A Practical Solution to the Reference Class Problem, forthcoming 109 Columbia Law Review (2009). The paper's abstract states:
    The "reference class problem" is a serious challenge to the use of statistical evidence that arguably arises every day in wide variety of cases, including toxic torts, property valuation, and even drug smuggling. At its core, it observes that statistical inferences depend critically on how people, events, or things are classified. As there is (purportedly) no principle for privileging certain categories over others, statistics become manipulable, undermining the very objectivity and certainty that make statistical evidence valuable and attractive to legal actors. In this paper, I propose a practical solution to the reference class problem by drawing on model selection theory in statistics. The solution has potentially wide-ranging and significant implications for statistics in the law. Not only does it remove another barrier to the use of statistics in legal decisionmaking, but it also suggests a concrete framework by which litigants can present, evaluate, and contest statistical evidence.

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Saturday, March 21, 2009

    Are People Just a Bundle of Traits: viz., Are Character Traits Reference Classes that Indicate How Individuals Likely Behave in Unknown Situations?

    I have long thought, often in an inchoate way, that thinking about the prediction of -- or inferences about -- human behavior provides important clues to the problem(s) of reference classes and, more broadly, to the manner in which experience teaches human beings -- or the manner in which human beings use experience -- to draw inferences about the world. Recently a member of my Advanced Evidence seminar asked the seminar members to read some things I wrote many years ago. On re-reading my own stuff, I concluded that what I said then was not entirely stupid, and I thought I would post that material on this blog for your consideration.

    The two batches of material are from Section 37 of my revision of the first volume of the fourth edition of Wigmore's multi-volume treatise on the American law of evidence. The first batch of material -- the much longer batch -- consists of footnote 8 of Section 37.6 of 1A Wigmore on Evidence (P. Tillers rev. 1983). The second extract consists of roughly two pages of text from Section 37.7.

    &&&

    Relative frequency, or frequentist, theories are supported by the general intuition that the known frequency of the association of two or more types of phenomena is a rational basis for making estimates of the probability of the association of those types of phenomena with each other in other cases. We have the intuition that the probability of such association in novel cases is normally a function, at least in part, of the relative frequency of such association of the various types of phenomena in cases about which we do have knowledge. Hence, we may rely on past frequency of association if we observe one phenomenon in a novel case and wish to estimate the probability of the occurrence of the other type of phenomenon.

    This general inclination to give credence to perceived regularities in the course of life and nature is the intuitive basis for the varieties of formal frequentist theories of probability. A primitive idea of relative frequency theory of probability is expressed in the Humean belief that something like a habit of thought, arising out of a perceived regularity of nature in known cases, gives some kind of reason to suppose it is probable that the sun will rise tomorrow.

    It is incontestable that our belief or perception of the existence of regularities in life and nature is a good basis on some occasions for our beliefs and estimates of the probability of a course of events that we have not been able to observe in some relatively "direct" fashion. But does this predisposition to give weight to some observed or perceived regularities mean that we should accept formal relative frequency theories of probabilities as an adequate general explanation of our interpretation of evidence? We think not — even though we believe that the human animal learns about conditions of existence through experience. To explain this view, we briefly examine more refined explanations of relative frequency theory.

    What is formal relative frequency theory? Burks has given this description: "[T]he frequency theory of empirical probability is the theory that atomic empirical probability statements should be analyzed into frequency probability statements and reasoned about by means of the calculus of frequency probability, which is an interpretation of the traditional calculus of probability." Burks, Chance, Cause, Reason: An Inquiry into the Nature of Scientific Evidence 148 (1977) (emphasis omitted). If the meaning of this definition is not entirely apparent, the more colloquial description given by John Maynard Keynes may prove to be more enlightening:

    "The essence of this theory can be stated in a few words. To say, that the probability of an event's having a certain characteristic is x/y is to mean that the event is one of a number of events, a proportion x/y of which have the characteristic in question; and the fact [original emphasis] that there is [original emphasis] such a series of events possessing this frequency in respect of the characteristic, is purely a matter of experience to be determined in the same manner as any other question of fact. That such series do exist happens to be a characteristic of the real world as we know it, and from this the practical importance of the calculation of probabilities is derived.

    "The two principal tenets . . . are . . . that probability is concerned with series or groups of events, and that all the requisite facts must be determined empirically." Keynes, Treatise on Probability 102-103 (1973) (reprint of first edition of 1921).

    The central tenet of frequentist theories of probability is the assertion that probability judgments are functions of the relative frequency with which two phenomena are associated with each other. More precisely stated, a relative frequency theory of probability supposes that the probability of event A, given B, is a function of the relative frequency with which A is known to occur when B occurs. The probability of A, given B, therefore, is a function of the ratio A/B. One theory of relative frequency, for example, provides that if we assume the existence of a "collective" of events, which we will call K, and if we wish to determine the relative frequency of a given type of event, which we will call E, in the collective K, we may say after a certain number of observations, which we will call n, of that type of event within the collective, that the relative frequency of E in K (P(E/K)) is equivalent to the limit that E/Kn approaches as n becomes large without bound. See Hacking, Logic of Statistical Inference 5 (1965) (description of refined version of theory developed by von Mises).

    In a more primitive form, the idea of relative frequency takes the form of direct (linear) extrapolation from existing observed regularities and recurrences. In this form, the idea of relative frequency may rely on the idea of direct linear extrapolation based upon simple enumeration: We count the frequency with which one event (or type of event) is associated with another event (or type of event). Call them A and B. To determine the probability of A, given B, we suppose (somewhat arbitrarily) that, absent other information, the observed relative frequency of A and B will hold in novel or observed instances and, given B, we therefore suppose that the likelihood of there also being an instance of A is equal to A/B.

    The proponents of relative frequency theory have advanced various, divergent formulas to describe the precise relationship between the relative frequency of A and B and a judgment of the probability of an instance of A, given B.

    We do not profess to grasp the mathematical considerations that have been urged in support of different mathematical formulas by which the lessons of known experience are extrapolated to novel situations. But we do believe, however presumptuously, that any such comprehension is immaterial for present purposes because we believe, on good authority, that such attempts to establish such formulas exhibit two related characteristics that illustrate why any rule of statistical inference is a bad candidate for general and universal validity.

    A rule of statistical inference is precisely a rule that works only if (1) it describes correctly the pattern of past experience, and (2) it describes a pattern that we know how to extrapolate to novel cases. Any proposed rule of statistical inference (from known experience to unknown cases), however, cannot by itself establish that a description of past experience is correct or accurate (in any meaningful sense) or that any pattern discernible in known cases should be extended in any particular fashion to unknown cases.

    Why is this so?

    The reason may be suggested by an assumed observation of the sequence "1, 2, . . . ?" What number follows 2? Our ability to answer depends, first, on our conviction that we have indeed observed "1" and "2" at the beginning of the series. (But how do we ever know that it is the beginning of the series that we have seen when indeed we have also observed other numbers at other times, on a digital printout, let us say?) But assume we know what we have seen. We must assume a certain principle of extrapolation. (Further observe that we may construct an innumerable variety of sequences into which "1, 2, . . . ?" would fit.)

    Similar difficulties arise if we assume that we have observed certain patterns of pairing of phenomena in the past, such as "a, b; a, c; a, b; a, d; a, b; . . .", and we wish to know P(b|a) (the probability of b, given a.) To extrapolate from the known cases, we must, first, be sure we have seen what we think we have seen. Second, we must say that some rule describes the pattern we have seen. But what rule? If we say that the rule is restricted to a description of what we have actually seen, we are not told how the rule should be extended to new cases. If, however, the rule purports to furnish a description of all pairings of a, b and not just those observed, some principle of generalization has been used. The difficulty, however, is that we may construct an infinite number of rules that are consistent with our known observations of a, b but that nonetheless describe different descriptions of the frequency of the pairing (and the distribution of the pairing) of a, b in the entire sequence (of both known and unknown cases). As we gather new observations, the difficulty merely reconstructs itself in a different form, one in which some series may be ruled out but one in which an infinite number of possible series remains. We need to make quite a few assumptions about patterns of divergence, convergence, uniformity, and stability before our aggregation of experience will in fact lead to a diminishing number of possibilities and an increased faith in particular series-descriptions. To be sure, in particular situations, particular sequences of events seem to make particular implications for the future (or for other situations) almost overpowering and practically irresistible. But as true as this point may be, it is largely immaterial, for we are here concerned about a theory of probability that explains all of our reasoning, and our attachment to particular sequences of connections, however powerful, is no basis for saying that we have discovered how to reason in all cases. Indeed, our problems of inference arise precisely because we seem to have no such powerful inferential sequence available to us in the case at hand, and we need to know what to do.

    But what of the principle of direct enumeration and direct extrapolation? Suppose we take a straightforward approach and assume that the probability of an instance of A in the future, once we know of B, is the ratio of the observed relative frequency of As to Bs in the past. Can we not use this principle, absent other information, on the general assumption that what has held true in the past (in known cases) will hold true in future or unknown cases in the same way? Contrary to all common sense, there are serious complications with this approach: "[T]he seemingly straightforward estimation of a probability order relation through induction by enumeration . . . requires . . . specification of the relevant set of past observations — the reference class of events; and an ability to recognize what is a confirming instance of the occurrence of an event in the data." Fine, Theories of Probability 112 (1973). See also Lonergan, Insight 302 (1978) (reprint of second edition of 1958) ("statistical laws presuppose some classifications of events").

    The problem of classification exhibits the fundamental weakness of any frequentist theory of probability. The power of any frequentist theory of probability ultimately depends on our ability to enumerate cases in appropriate and meaningful ways, and our power to enumerate, of course, depends on our ability to recognize whether particular events belong within some groups, class, or type of event of which we wish to make some sort of enumeration. This act of classification, however, is not always a self-executing act whose legitimacy cannot be questioned. We find, thus, that the act of classification may depend upon the perception of an analogy between one event and another, which, sometimes, leads us to say that both events belong within some common class of events. See de Finetti, Probability, Induction and Statistics: The Art of Guessing 154 (1972) ("The special case of statistics, according to our interpretation, differs from those illustrated in the preceding examples only in that the observable events E1, . . ., Eq instead of being diversified are analogous, or (according to a terminology that I regard as inadmissible) identical" (original emphasis)).

    It is true that statistical enumeration does in some cases serve as a powerful predictive tool. But cf. Northrop, The Logic of the Sciences and the Humanities 115 (1947) ("David Hume, who was the first Western thinker perhaps to fully realize the exact character of purely empirically given knowledge, pointed out that such knowledge exhibits no necessary connections. This means that a science which restricts itself to directly observable entities and relations automatically loses predictive power. The science tends, even when deductively formulated, to be merely descriptive and to accomplish little more so far as prediction is concerned than to express the hope that the sensed relations holding between the entities of one's subject matter today will recur tomorrow. This is an excessively weak and deductively empty type of predictive power. Little can be deduced from mere subjective psychological hope"). The genuine power of statistical reasoning within some domains neither demonstrates that statistical reasoning, based on the principle of enumeration of like or "identical" cases, is powerful within all domains nor that all inferential reasoning is based on it. The phenomenon of "stable measurement" accounts for the successes of statistical reasoning. Cf. de Finetti, Probability, Induction and Statistics: The Art of Guessing 145 (1972). See also Hempel, Aspects of Scientific Explanation 386 (1965) ("The mathematical theory of statistical probability is intended to provide a theoretical account of the statistical aspects of repeatable processes of a certain kind which are referred to as random processes or random experiments"). In many cases, however, such stable measurement does not exist and cannot be made to exist unless we choose to be arbitrary. In these cases, the value of statistical measurement is uncertain. (The difficulty of determining whether an event is to be regarded as being of a particular type has helped to inspire some mathematical theories about "fuzzy sets" in which this problem is explicitly acknowledged. See Zadeh, Fuzzy Sets, 8 Information & Control 388 (1965). Our argument, however, makes plain that the mathematical concept of a fuzzy set will not resolve all of our difficulties. Here, no less than elsewhere, the facts will not always speak for themselves, and our interpretive rules may be seen as being somehow prior to the data being observed and classified.

    The difficulty suggested by the problem of measurement and classification may be described in another way. Suppose that sets of phenomena are infinitely rich and diverse and that we "partition" the phenomena in different ways when we use different names and classifications to denote what we observe in a given set of complex phenomena. What reason do we have to suppose that the partitioning of a particular set of phenomena has been made in a useful way? Here again, sometimes it may happen that the phenomena in question will seem almost to partition or classify themselves in a useful and convincing way, but in other cases this will not happen and then we are again faced with the problem of arbitrary classifications and partitions. Cf. de Finetti, Probability, Induction and Statistics: The Art of Guessing 155 (1972) ("The validity of a given property (such as some form of the law of numbers) never depends on similarities or on any external features of the events concerned, but only on the probability scheme accepted for them. The external features are relevant only in the role they play in determining our opinion about the probabilities"). If we suppose that we can always avoid having to decide which particular partition of a number of possible partitions is "correct," we are quite mistaken, because we will find that in many cases statistical extrapolations from different partitions of the data will lead to quite different conclusions. This may be illustrated by the infamous problem involving a Swedish pilgrim to Lourdes. What is the probability that the pilgrim is Catholic if 95 percent of Swedes are not Catholic and if 95 percent of pilgrims to Lourdes are Catholic? See, e.g., Ayer, Probability and Evidence 51-52 (1972). This problem has been called the problem of intersecting or overlapping reference classes. The discussions of this problem and similar problems show that statistical reasoning alone is here helpless and leads to contradiction.

    Carl Hempel describes the problem of intersecting classes thus:

    "[The ambiguity of statistical explanation] derives from the fact that a given individual event . . . will often be obtainable by random selection from any one of several `reference classes' . . ., with respect to which the kind of occurrence . . . instantiated by the given event has very different statistical probabilities. Hence, for a proposed probabilistic explanation with true explanans which confers near-certainty upon a particular event, there will often exist a rival argument of the same probabilistic form which confers near-certainty upon the nonoccurrence of the same event" (Aspects of Scientific Explanation 394-395 (1965)).

    The problem of intersecting classes may be stated in a somewhat different form: If probability statements are a product of frequency statements, it is logically incompatible with the axioms of probability theory to state, for example, both that the probability of X is .95 and to state that the probability of not-X is .95, for P(X) = 1 — P(not-X). Hence, some modification of the implications of conflicting frequency observations must be made on some basis not generated by relative frequencies alone. See Hempel, Aspects of Scientific Explanation 72-73 (1965).

    Carnap tried to resolve the paradox of conflicting reference classes by two devices. First, he devised a theory of what has been called tautological probability, which defines probability as the distribution of a selected characteristic or event over some chosen reference class. There is then no conflict because a separate reference class has, by definition, a separate probability distribution. This solution, though logically permissible, achieves a Pyrrhic victory because the notion of probability is made unusable in reality. Second, Carnap advocated a principle of "total evidence," by which he meant, apparently, that one would rely on all available information, which in this context presumably means that one should choose a reference class that includes (by definition) only "Swedish pilgrims to Lourdes" and not "Swedes" or "pilgrims to Lourdes." This latter solution, called the "requirement of maximal specificity," is also impracticable, for reasons we cannot review here. See, e.g., Hempel, Aspects of Scientific Explanation 394-402 (1965).

    In summary, the choice of a classification of events, the selection of events as falling within any such classification, and, furthermore, the selection of a principle or rule by which the relative frequency of such events is extended to cases of interest — none of these choices is self-executing. Rather, each choice requires the exercise of human judgment, by which means some pattern is imposed on human experience. Accordingly, it is untenable to say that experience alone is the basis for inference, and accordingly, it seems clear that the principle of relative frequency — statistical frequency — is not sufficient to explain the interpretation of evidence. (Affirmatively, we assert that the principle of relative frequency achieves power only if the organizing activities of the observer are recognized to be indispensable.)

    The foregoing objections constitute an argument against simple empiricism. The empiricist notion is that experience speaks for itself by exhibiting a certain regularity, but the intrinsic complexity of phenomena prevents any sort of mechanical extrapolation from experience. Any phenomenon can be partitioned, classified, or "experienced" in innumerable different ways. Thus, it follows that of three experiences, all three will be the same or similar in certain respects — in an infiite number of respects — that all three will be different in certain respects — in an infinite number of respects — and that what has just been said will also hold for any two of the three experiences. It is also true that these differentiating and nondifferentiating characteristics of experiences appear themselves as a composite of innumerable characteristics and that what has just been said of "experiences" also holds true for the characteristics by which we attempt to differentiate and relate different experiences. Furthermore, it is also true that experience may well disclose an infinite number of regularities and patterns since experience (we suppose) may be decomposed in an infinite variety of ways. If these assumptions are correct, it follows that experiences do not of themselves exhibit or establish their differences and similarities and that experiences of themselves do not even tell us whether we are seeing the same thing (the same connection) we saw before. In other words, an experience so diverse cannot determine whether we have experienced any regularity in the course of nature. If the conclusion seems absurd, it is because strict empiricism is absurd. Empiricism avoids these difficulties only by assuming the legitimacy of classification of experience, by assuming we know how to classify, or by assuming that nature classifies itself.

    Different responses are possible to these difficulties of frequentist theories of probability. One response is subjective probability theory, which essentially abandons the attempt to base probability in the objective features of the world. We prefer a different response. We do not think the inadequacies of frequentist theories mandate a flight into solipsism. What is required is our recognition that our inferences from evidence always involve some sort of "contribution" by the factfinder, by which experience is organized into certain patterns that are not themselves inexorably given by experience. There is something like a "web of belief," by which we organize, wittingly and unwittingly, our experience of regularity. Thus, there now exist what are called "presupposition" theories of probability, which are, in part, what their name implies. See Burks, Chance, Cause, Reason: An Inquiry into the Nature of Scientific Evidence 647-650 (1977) (summarizing theories). These theories recognize that any theory of relative frequency or induction on the basis of relative frequency requires or presupposes (by its very existence, perhaps) some principle of enumeration that guides the methods of our enumeration of the data. We might well say that the development of such a principle of enumeration is the precise object of a theory of statistics or induction. Without such a principle, our counting is pointless and aimless.

    There are other difficulties with relative frequency theory, but we view these as being of secondary importance. See, e.g., Ayer, Probability and Evidence 51-53 (1972) (making a distinction between statistical statements and judgments of credibility that deal with what has been called the "unique case" situation; we choose to relate this difficulty to the problem of in tersecting classes or, what is the same thing, to the question of making appropriate partitions of infinitely varied data; see discussion above).

    In some cases, we may not know the genesis of our principle or procedure of counting, but we still know that the human organism does add such a principle of enumeration (expressly or implicitly) when it does choose to count and to rely on such counting. Hence, we may also assert that in some cases we have conceptual presuppositions (whatever their source) that may tell us (expressly or implicitly) that counting in a mechanistic fashion (by any rule) is either inappropriate or insufficient. If so, it is not at all odd to suppose that in some cases an observer is simply incapable of organizing the evidence before him clearly enough even to imagine the possibility of counting, and it is not at all absurd to suppose that some evidence will never in principle be understood by the observer to be sufficiently "atomistic" to be capable of being counted. In such a case, it is not demonstrably irrational to suppose that forced counting is a less accurate and reliable method for performing the task in question than some other less "precise" method of inference from evidence. We need not indulge in a metaphysical hypothesis that the observer and factfinder must be wrong in this supposition, and we need not imagine that we can improve his factfinding skill by making him more "rational" in the sense understood by one who believes in the universal validity and applicability of relative frequency theory. Relative frequency theory is a pretty model that will not work in some places.

    &&&

    [L.J.] Cohen's theory deepens and advances our understanding of the complex way in which the mind of the observer, through generalizations and the like, uses beliefs and principles to evaluate the probative force of a given piece of evidence, and he shows us that there are problematic features to the observer's use of generalizations. However, in our view, Cohen does not take us far enough, either qualitatively or quantitatively. Qualitatively, he does not take us far enough because, all provisos considered, he still takes the view that the interpretive conceptual principle that speaks to the probative force of a piece of evidence in essence still amounts to a statement that describes (within its appropriate domain) certain events that occur with a certain frequency relative to other events. However, there are conceptual interpretive structures that, though speaking to the probative force of a piece of evidence, take a quite different form. The term "generalization" implies that the beliefs and theories and concepts of the observer always amount to a generalized description of the course of nature that constitutes an extrapolation from regularities noted by the observer in a limited number of instances. However, "experience" can work in quite different ways and may lead to the formation of conceptual and interpretive systems that cannot easily be described as statements that describe the relative frequency of various types of events under various conditions. We do not know what sort of name to give to such conceptual and interpretive structures, but we may illustrate what we mean by an example. This example suggests that a different kind of inferential process may often apply to the assessment of the probable course of a person's behavior.

    Ordinary "generalizations," like scientific theories, may have a complex logical structure and may involve complex, though largely implicit, logical operations and calculations. Consider, for example, a defamation action in which the plaintiff alleges that the defendant called him a "son of a bitch" during a radio broadcast. At the trial, the plaintiff offers into evidence a tape of that program. The tape, unfortunately, is inaudible at certain points, and the jury can only hear the words "You are a son of a ****." The question, then, is whether the defendant used the word "bitch" or some other word, such as "gun." Now suppose that the jury has the hunch that the defendant said "bitch." Does it have this hunch simply because it adheres to certain implicit generalizations about the frequency with which the word "bitch" follows the words "You are a son of a"? No matter what kind of a generalization of this sort we look for, we are likely to oversimplify how the jury is making its calculations. It is quite likely that the jury's hunch in large part rests on intuitive but complex assessments of the rules that govern our modes of speech. Thus, for example, the jury may rule out the possibility that the defendant said, "You are a son of a water," not because it has observed that "water" is rarely used by persons in such a context but rather because it has decided that such a phrase has no sensible meaning and that no rational person, adept in the English language, would ordinarily say such a meaningless thing. The jury thus rules out this possibility on some sort of logical ground. Similarly, the jury is likely to rule out other candidates such as "ape," not only because "ape" is rarely used in such a sentence, but also because its use there would not be expected because of the ungrammatical construction that would result ("son of a ape"). In short, intuitive reasoning about language usage and its conventions reduces the likely candidates here to "gun" and "bitch," (Of course the jury's reasoning and hunches become more complicated if the antecedent conversation and antecedent statements by the defendant are taken into account. Here again, however, the jury is likely in part to rely on various implicit rules about appropriate conventional usage of language, and these rules of usage, like explicitly formulated rules of grammar and usage, are rules that the jury somehow knows it must interpret in order to predict what specific usage is likely to occur in a situation the jury has not previously encountered. Thus, if the antecedent conversation had related to family trees and royalty, the jury may infer from this conversation that the word chosen was "Mountbatten." And the jury might reach this conclusion even if it had never before heard talk about the Mountbatten family or about royal lineage and thus had had no previous experience with the phrase "son of a Mountbatten." It is a rule or principle of usage, implicitly but logically extended to a novel situation, that informs the jury that "Mountbatten" is the type of word (like "Smith") that may be used in this sentence.)

    Though we may be incapable of formulating all the principles and operations that lead a judge or jury to the hunch that the word "Mountbatten" or "bitch" was probably used (since, for example, the tone of voice used by the speaker, as well as the pacing of the words, has a bearing on the question), it does not follow that the judge (for example) lacks all capacity to investigate and explicate the bases of that inference of his; and it is also possible that such self-conscious investigation, though necessarily partial, may yet lead the judge both to revise his hunch and also to have more confidence in the reliability of whatever hunch he eventually does have.

    It is important to recognize that the use of language does not present an isolated or unusual example of the complex character of the inferential processes that are normally involved in the assessment of factual issues. Brief reflection on the types of factual issues submitted for adjudication quickly reveals that many judicial factfinding efforts — if not most of them — involve attempts to assess the probable behavior of some person or persons within a social context. As with language, social behavior is governed by complex systems of unspoken complexes of rules that, while not inexorably determining how a human being behaves within a particular social context, do suggest how persons normally behave and may be expected to behave within such a context, and it is surely the case that we at least implicitly refer to such complexes of rules and principles in trying to determine what the actor probably did.

    These rules, no less than the rules of grammar and of language usage, are surely quite complex and have a kind of logic of their own; they do not amount to simple generalizations, based on prior experience with a similar situation, about what sort of behavior usually occurs in a particular setting. As in the case of language, introspection and reflection, while presently incapable of any exhaustive description of the social and interpersonal rules people follow, may help us see whether the peculiar features of a particular social setting are likely to have affected the behavior of an actor given the sort of "logic" of social relations that he, like ourselves, tends to follow.

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Thursday, March 19, 2009

    Law Reviews and Legal Scholarship in the Age of Cyberspace

    It would have been hard to imagine just a decade ago: major U.S. newspapers are going out of business. (Quixotically enough, I recently renewed my subscription to hard copies of the New York Times.) The same fate has not yet befallen U.S. law reviews.

    But most American law reviews are not, and perhaps never were, subject to normal market forces; most of these student-edited law journals were and are subsidized by law schools.

    Subscriptions to major law reviews have fallen dramatically in the last couple of decades. Is the demise of student-edited law reviews at hand?

    Well, even if we ignore the nebulousness of the notion of "demise," it's not yet clear that Armageddon for law reviews is at hand. This is because law schools have non-economic reasons for wanting to keep law journals alive.

    It may be true -- though demonstrating this would be tricky -- that most "major" American student-edited law reviews are kept alive in significant part because "major" law schools want to maintain some control over access to the halls of legal academe and over the kinds of scholarship that secure access to U.S. legal academe. But there are signs that the gatekeeper role of these law reviews is on the wane.

    That's probably a good thing.

    The market, she is tricky, fickle, and often downright stupid. But the market is also often relatively democratic and open to innovation.

    In the age of cyberspace budding legal scholars have some serious alternatives to student-edited "major" law reviews.

    It is true that law schools will very probably still use "major" hard-copy student-edited law reviews as gatekeepers. But cyberspace and other developments are gradually creating alternatives to "major" law schools themselves. As California's Bernard Witkin demonstrated decades ago, such alternatives always existed. But in the age of cyberspace the prospects for market-oriented legal scholarship have grown and multiplied.

    &&&

    I confess that personal history motivates this post. Decades ago, I swore not to submit my stuff to "major" American law reviews. I departed from my populist anti-establishmentarian line generally only when a law journal invited me to submit a paper. Otherwise I have published in other venues. I took this anti-establishmentarian tack when, shortly after graduation, I tried to publish a study of Hegel's theory of the "duty to die for the state." I submitted my paper to about five "major" law reviews. They rejected my paper (but, in fairness to them, usually only by close votes).

    I later realized I was literally ahead of my time: Hegel was not yet in vogue in American law schools. Had I tried publishing the paper a couple of decades later, I would have met with success. But by then I had completely repudiated Hegel and I had little taste for talking about things Hegelian. (The rejected paper was an excellent piece of work. [I concluded that Hegel's argument for the alleged duty to die for the state fails.])
    This experience led me to swear off law reviews. I instead worked at redoing part of Wigmore's treatise.

    Of course, by swearing off law reviews (for the most part) I figuratively shot myself in my figurative academic foot. But I don't regret what I did. I think my scholarship was more interesting as a result. I discovered, to my pleasure if not entirely to my surprise, that there are lots of inquisitive, creative, serious, and thoughtful people out there in the legal profession and in the wider world. Conclusion: publishing for the "market" and for the "world" has its compensations, very substantial compensations.

    Postscript: Bernard Witkin's model of legal scholarship is not the model to which I aspire. But that's another question. My point here is that Witkin succeeded in doing legal scholarship on his own.

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    The High Probability of Very Improbable Events

    "'Everything we see has about a zero probability,' [Peter H.] Westfall said." Carl Bialik, "The Crash Calculations," THE NUMBERS GUY (March 3, 2008).

    "'With a large enough sample, any outrageous thing is apt to happen.'" Gina Kolata, "1-in-a-Trillion Coincidence, You Say? Not Really, Experts Find," NYTimes (Feb. 27, 1990) (quoting statisticians Persi Diaconis and Frederick Mosteller).

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Sunday, March 15, 2009

    MarshalPlan 2.3, a System for Marshaling Evidence in Legal Settings such as Trials and Pretrial Investigations

    I have slightly updated and posted my evidence marshaling "stacks" (software) for the software MarshalPlan. Below please find general information about MarshalPlan and about how to download the software.

    &&&

    Years ago David Schum and I developed the notion of an evidence marshaling system. We laid out the underlying theory of this evidence marshaling system in A Theory of Preliminary Fact Investigation. We developed a kind of computer embodiment, or computer-based expression, of our idea of an evidence marshaling system. Eventually we decided to call our system "MarshalPlan."

    About one year ago I released MarshalPlan 2.2. This moniker -- MarshalPlan 2.2 (now 2.3) -- amounted to a bit of self-mockery: MarshalPlan 2.x is not a software "prototype." Far from it! However, MarshalPlan 2.2 and 2.3 are more than scratchings on a page that state in words (text) how a MarshalPlan application might work. MarshalPlan 2.3 is a software application based on the user-friendly programming language Revolution Enterprise(tm). This application -- MarshalPlan 2.3 -- illustrates -- with images, fields, buttons (links), and so on -- how a computer program to support the marshaling and assessment of evidence in preparation for possible trials and also for the conduct of trials, might work.

    &&&

    To retrieve MarshalPlan 2.3 click on this link. Download all of the Revolution stacks into a single folder on your computer. These stacks all have the suffix "rev". To make these stacks run properly you need a "Revolution Player." To get this free player go here and download the version of the player (either Windows or Mac OSX) that you need. Then drag-drop the "Network.rev" icon onto the "Revolution Player" icon or open the Revolution Player icon and then open the Network.rev stack, or file. You should be in business now; the buttons, or links, in the various stacks should allow you to navigate between the stacks as well as within the stacks. (However, it is possible you will have to drag-drop all of the stacks onto the Revolution Player icon if you wish to navigate between the stacks. Please let me know if this turns out to be the case.)

    &&&

    SOME VERY IMPORTANT CAVEATS: There are numerous things wrong with the software application that you will retrieve by clicking on the link or links below, and the application that you will retrieve has numerous gaps and defects, including the following:

    1. In the application itself there is very, very, very little textual explanation of the theory behind the strategies that are embedded in MarshalPlan 2.3.
    To find that theory and those explanations you will have to (i) read the article I mentioned earlier, A Theory of Preliminary Fact Investigation, and (ii) wander about my personal web site. If you want a really comprehensive explanation of MarshalPlan, you will have to invite me to give a leisurely talk (preferably on a tropical island or some other attractive venue).
    2. Some buttons and links don't work. When that happens, try other buttons and links.

    3. Some important stacks are entirely missing. E.g., the "Narratives" stack. The most important missing stacks are those having to do with the development of evidential argument from evidence to factual propositions and with the assessment of the probative value of the evidence. For a discussion of the methods that might be used for this purpose, see Special Issue on Graphic and Visual Representations of Evidence and Inference in Legal Settings, 6 Law, Probability and Risk Nos. 1-4 (Oxford University Press, 2007).

    4. MarshalPlan 2.3 is not set up to be linked to a database. This is a most serious deficiency. But -- in my defense -- I repeat: MarshalPlan 2.3 is NOT a software prototype. It is, rather, an elaborate visual illustration of some of the directions that development of software for marshaling evidence in legal settings should take.

    &&&

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Saturday, March 14, 2009

    The Near-Term Prospects for fMRI Lie Detection

    2 Margaret A. Boden, MIND AS MACHINE 1227-1228 (Oxford: Clarendon Press 2006):

    Linking rats or monkeys to robots is a major exercise, not undertaken lightly. But brain-scanning studies on humans are less tricky, and are multiplying merrily. ... As for the professional journals, by the time you read this book they will have carried thousands of PET/fMRI reports.

    Their theoretical significance, however, is debatable. Brain imaging has even been dubbed "a neo-phrenological fad", because of the difficulty of interpreting it in terms of psychological functions (Uttal 2001). There are four main problems.

    [snip, snip]

    And fourth, one can't sensibly suggest just what's being done by the high activity (even if one knew it was excitatory activity) without a theory at the cognitive/psychological level, specifying just what computations might be involved when the thought in question occurs. Usually, no such theory is available. ... In short, most brain imaging is an a-theoretical fishing expedition: more natural history than science. As Gazzinga had put it, before this new 'industry' burgeoned, neuroscience needs cognitive science.

    Responsible researchers know all this, of course, and are careful. But the irresponsible ones--and a fortiori the journalists--seemingly don't, and aren't.

    FINIS

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Friday, March 13, 2009

    The Sophistication of (Some) Common Sense

    In thinking about juries, jurors, lay knowledge, and lay participation in the legal process, it is worth thinking about the implications of insights such as the following:
    Men are in a more difficult intellectual position than Life robots. We don't know the fundamental physics of our world, and we can't even be sure that its fundamental physics is describable in finite terms. Even if we knew the physical laws, they seem to preclude precise knowledge of an initial state and precise calculation of its future both for quantum mechanical reasons and because the continuous functions needed to represent fields seem to involve an infinite amount of information.

    This example suggests that much of human mental structure is not an accident of evolution or even of the physics of our world, but is required for successful problem solving behavior and must be designed into or evolved by any system that exhibits such behavior.

    John McCarthy, "Ascribing Mental Qualities to Machines" (1979)" (with updates by author here & there)

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law

    Tuesday, March 10, 2009

    The Riddle of Non-Individualized Forensic Expertise

    Well, that -- the topic in the header -- is a real mouthful. It is nevertheless the problem I want to say a few words about.

    In my last post I quoted the Texas Court of Criminal Appeals. That court said:

    Judge Womack contends that our exclusion of expert conclusions concerning the truthfulness of allegations is illogical. He isolates three different types of statements that an expert might make: (1) children who are fantasizing or being manipulated behave in certain ways, (2) this child did not behave in those ways, and (3) this child was not fantasizing or being manipulated. He claims that (3) necessarily follows from (1) and (2), and, because we permit expert statements of the type (1) and (2) variety, we must necessarily permit conclusions of type (3). But, for his argument to work, Judge Womack must assume that the expert testifies that all children who are fantasizing or being manipulated behave in certain ways. If only some exhibit the behaviors in question then a conclusion that a particular child is not fantasizing or being manipulated because he does not exhibit the behaviors does not necessarily follow. But, given the imprecision of psychological science and the variability of human nature, no competent and honest expert could make the global statement necessary to satisfy Judge Womack's syllogism. At most, statements of type (1) and (2) would provide some inductive support for a conclusion of type (3). But an expert opinion of type (3) would also be supported by personal observations and lay knowledge of human behavior. The latter is clearly the exclusive province of the jury, and the former, along with statements of type (1) and (2) can be imparted to the jury by the expert. The jury can then make the inferences necessary to determine whether a type (3) conclusion is warranted without the expert commenting on that issue.
    Here's the riddle: If the expert's expertise is relevant to the issue at hand (e.g., "Was this particular eyewitness identification accurate or inacurate", "Does this particular witness suffer from delusions?", "Does the syndrome evidence show that this particular child probably delayed reporting because of embarrassment rather than for another reason?", and so on), why isn't the expert not only permitted but required to give an opinion about the behavior of the specific individual? Is it because the expert has no expertise about the specific individual? Well, no that cannot be -- for then the expert's expertise would not be relevant to the actual issue at hand. Is it because the expert would invade the province of the jury by giving an opinion? But courts have already crossed that imaginary Rubicon by allowing the expert to testify, haven't they? Well, then, is it because the expert doesn't know all the evidence and facts about the individual that the jury knows? Well, that's not going to explain things, is it, if the expert has sat through the entire trial and knows everything (and more) than the jury knows. So what is the explanation? We can't very well say (as the Texas court seems to suggest) that it's because the jury knows better than the expert how to combine different kinds of information (or, as some would put it, different reference classes). If that were the explanation, it would require the premise that the experimental evidence of the expert failed to take into account relevant variables or, stated differently, that the expert's expert knowledge does not speak to the individual, particularized issue at hand.

    Isn't it the case that the position that courts take -- let the expert testify in general terms (about, e.g., factors that affect the accuracy of eyewitness identification in general) but not in particular terms (e.g., about the accuracy of this witness' identification) -- is an unprincipled compromise and muddle, one that straddles the fear of doing without helpful information and the fear of allowing the decision in a case depend on unhelpful or irrelevant information?

    the dynamic evidence page

    coming soon: the law of evidence on Spindle Law