|
|
FORENSIC SCIENCE EVIDENCE IN LITIGATION
101
standards from the beginning,62 because scientific groundwork for DNA
analysis had been laid outside the context of law enforcement. The National
Institutes of Health (NIH) and other respected institutions funded and
conducted extensive basic research, followed by applied research. Serious
studies on DNA analysis preceded the establishment and implementation
of “individualization” criteria and parameters for assessing the probative
value of claims of individualization. This history stands in sharp contrast
to the history of research involving most other forensic science disciplines,
which have not benefitted from extensive basic research, clinical applica-
tions, federal oversight, vast financial support from the private sector for
applied research, and national standards for quality assurance and quality
control. The goal is not to hold other disciplines to DNA’s high standards
in all respects; after all, it is unlikely that most other current forensic
methods will ever produce evidence as discriminating as DNA. However,
using Daubert as a guide, the least that the courts should insist upon from
any forensic discipline is certainty that practitioners in the field adhere to
enforceable standards, ensuring that any and all scientific testimony or
evidence admitted is not only relevant, but reliable.
Judicial Dispositions of Questions Relating to Drug Identification
Over the years, there have been countless instances in which trial judges
have assessed the admissibility of expert testimony relating to drug analy-
ses, either sua sponte or pursuant to objections raised by defense counsel.
Because trial court decisions in these matters often are resolved without
published written opinions and with no challenges on appeal, there is no
sure way to know how often trial judges deny the admissibility of the evi-
dence. Trial judges may sometimes sustain challenges to the admissibility
of expert testimony, especially in instances where the defense can show
defects in the foundational laboratory reports.63 But there are very few
such reported cases.
In addition to alleged defects in laboratory reports and sampling pro-
cedures, trial courts routinely consider whether experts possess the neces-
sary qualifications to testify and, more generally, whether expert testimony
is sufficiently reliable to be admitted under Daubert and Federal Rule of
Evidence 702. However, in published opinions addressing expert testimony
based on drug identification, federal appellate courts rarely reverse trial
62 See supra text accompanying note 54; see also Gov’t of V.I. v. Byers, 941 F. Supp. 513
(D.V.I. 1996); United States v. Jakobetz, 747 F. Supp. 250 (D. Vt. 1990), aff’d, 955 F.2d 786
(2d Cir. 1992).
63 See, e.g., United States v. Diaz, 2006 WL 3512032 (N.D. Cal. 2006).
102
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
court decisions rejecting Daubert challenges.64 Why? First, as noted above,
in cases where the evidence is excluded at trial, no appeal will be taken.
Second, the scientific methodology supporting many drug tests is sound.
This means that, regardless of the standard of review, most decisions by
trial courts will withstand scrutiny. Finally, courts of appeals owe great
deference to trial court judgments on questions relating to the admission
of evidence.65
The importance of the limited standard of review was clearly explained
in United States v. Brown:66
Immersed in the case as it unfolds, a district court is more familiar with the
procedural and factual details and is in a better position to decide Daubert
issues. The rules relating to Daubert issues are not precisely calibrated
and must be applied in case-specific evidentiary circumstances that often
defy generalization. And we don’t want to denigrate the importance of the
trial and encourage appeals of rulings relating to the testimony of expert
witnesses. All of this explains why the task of evaluating the reliability of
expert testimony is uniquely entrusted to the district court under Daubert,
and why we give the district court considerable leeway in the execution
of its duty. That is true whether the district court admits or excludes ex-
pert testimony. Joiner, 522 U.S. at 141-42 (“A court of appeals applying
‘abuse-of-discretion’ review to [Daubert] rulings may not categorically dis-
tinguish between rulings allowing expert testimony and rulings disallowing
it.”). And it is true where the Daubert issue is outcome determinative.67
Judicial Dispositions of Questions Relating to Fingerprint Analyses
Over the years, the courts have admitted fingerprint evidence, even
though this evidence has “made its way into the courtroom without empiri-
cal validation of the underlying theory and/or its particular application.”68
The courts sometimes appear to assume that fingerprint evidence is irrefut-
able. For example, in United States v. Crisp, the court noted that “[w]hile
the principles underlying fingerprint identification have not attained the
64 See, e.g., United States v. Moreland, 437 F.3d 424, 430-31 (4th Cir. 2006), cert. denied,
547 U.S. 1142 (2006); United States v. Scalia, 993 F.2d 984, 988-90 (1st Cir. 1993).
65 See, e.g., United States v. Gaskin, 364 F.3d 438, 460 n.8 (2d Cir. 2004) (holding that
“when a party questions whether sound scientific methodology provides a basis for an expert
opinion, it may move to preclude the admission of the opinion” under Daubert; however,
when a defendant makes no such motion and instead stipulates to the admissibility of the
expert opinion, “he cannot complain on appeal that the opinion lacks foundation”).
66 415 F.3d 1266 (11th Cir. 2005).
67 Ibid., pp. 1265-66 (alteration in original) (internal quotation marks, other internal cita-
tions omitted).
68 M.A. Berger. Procedural paradigms for applying the Daubert test. 78 Minn. L. Rev.
1345, 1354 (1994).
FORENSIC SCIENCE EVIDENCE IN LITIGATION
103
status of scientific law, they nonetheless bear the imprimatur of a strong
general acceptance, not only in the expert community, but in the courts as
well.”69 The court went on to say:
[E]ven if we had a more concrete cause for concern as to the reliability of
fingerprint identification, the Supreme Court emphasized in Daubert that
“[v]igorous cross-examination, presentation of contrary evidence, and
careful instruction on the burden of proof are the traditional and appropri-
ate means of attacking shaky but admissible evidence.” Daubert, 509 U.S.
at 596. Ultimately, we conclude that while further research into fingerprint
analysis would be welcome, “to postpone present in-court utilization of
this bedrock forensic identifier pending such research would be to make
the best the enemy of the good.”70
Opinions of this sort have drawn sharp criticism:
[M]any fingerprint decisions of recent years . . . display a remarkable lack
of understanding of certain basic principles of the scientific method. Court
after court, for example, [has] repeated the statement that fingerprinting
met the Daubert testing criterion by virtue of having been tested by the
adversarial process over the last one-hundred years. This silly statement is
a product of courts’ perception of the incomprehensibility of actually limit-
ing or excluding fingerprint evidence. Such a prospect stilled their critical
faculties. It also transformed their admissibility standard into a Daubert-
permissive one, at least for that subcategory of expertise.71
This is a telling critique, especially when one compares the judicial decisions
that have pursued rigorous scrutiny of DNA typing with the decisions that
have applied less stringent standards of review in cases involving fingerprint
evidence.
In holding that fingerprint evidence satisfied Daubert’s reliability
and relevancy standards for admissibility, the Fourth Circuit’s decision
in Crisp noted approvingly that “the Seventh Circuit [in United States
v. Havvard, 260 F.3d 597 (7th Cir. 2001)] determined that Daubert’s
‘known error rate’ factor was satisfied because the expert in Havvard
had testified that the error rate for fingerprint comparison was ‘essentially
zero.’”72 This statement appears to overstate the expert’s testimony in
Havvard, and gives fuel to the misconception that the forensic discipline
69 324 F.3d 261, 268 (4th Cir. 2003).
70 Ibid., pp. 269-70 (second alteration in original) (other internal citation omitted).
71 1 Faigman et al., op. cit., supra note 1, § 1:1, p. 4; see also J.J. Koehler. Fingerprint er-
ror rates and proficiency tests: What they are and why they matter. 59 Hastings L.J. 1077
(2008).
72 324 F.3d at 269 (quoting Havvard, 260 F.3d at 599).
104
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
of fingerprinting is infallible. The Havvard opinion actually described the
expert’s testimony as follows:
[The expert] testified that the error rate for fingerprint comparison is
essentially zero. Though conceding that a small margin of error exists
because of differences in individual examiners, he opined that this risk is
minimized because print identifications are typically confirmed through
peer review. [The expert] did acknowledge that fingerprint examiners have
not adopted a single standard for determining when a fragmentary latent
fingerprint is sufficient to permit a comparison, but he suggested that the
unique nature of fingerprints is counterintuitive to the establishment of
such a standard and that through experience each examiner develops a
comfort level for deciding how much of a fragmentary print is necessary
to permit a comparison.73
This description of the expert’s equivocal testimony calls into question any
claim that fingerprint evidence is infallible.
The decision in Crisp also pointed out that “[f]ingerprint identification
has been admissible as reliable evidence in criminal trials in this country
since at least 1911.”74 The court, however, pointed to no studies supporting
the reliability of fingerprint evidence. When forensic DNA first appeared, it
was sometimes called “DNA fingerprinting” to suggest that it was as reli-
able as fingerprinting, which was then viewed as the premier identification
science and one that consistently produced irrefutable results. During the
effort to validate DNA evidence for courtroom use, however, it became
apparent that assumptions about fingerprint evidence had been reached
without the scientific scrutiny being accorded DNA. When the Supreme
Court decided Daubert in 1993, with its emphasis on validation, legal com-
mentators turned their attention to fingerprinting and began questioning
whether experts could match and attribute fingerprints with a zero error
rate as the FBI expert claimed in Havvard, and whether experts should be
allowed to testify and make these claims in the absence of confirmatory
studies. As noted above, most of these challenges have thus far failed, but
the questions persist.
The 2004 Brandon Mayfield case refueled the debate over fingerprint
evidence. The chronology of events in the Mayfield case is as follows:
73 Havvard, 260 F.3d at 599. The Havvard decision is sharply criticized by 1 Faigman et al.,
op. cit., supra note 1, § 1:30, pp. 86-89.
74 Crisp, 324 F.3d at 266. The decision cites a number of other legal references, includ-
ing, inter alia: People v. Jennings, 96 N.E. 1077 (1911); J.L. Mnookin. Fingerprint evidence
in an age of DNA profiling. 67 Brook. L. Rev. 13 (2001) (discussing history of fingerprint
identification evidence).
FORENSIC SCIENCE EVIDENCE IN LITIGATION
105
March 11, 2004: Terrorists detonate bombs on a number of trains in
Madrid, Spain, killing approximately 191 people, and injuring thousands
more, including a number of United States citizens.
May 6, 2004: Brandon Bieri Mayfield, a 37-year-old civil and immigration
lawyer, practicing in Portland, Oregon, is arrested as a material witness
with respect to a federal grand jury’s investigation into that bombing. An
affidavit signed by FBI Special Agent Richard K. Werder, submitted in sup-
port of the government’s application for the material witness arrest war-
rant, [avers] that Mayfield’s fingerprint has been found on a bag in Spain
containing detonation devices similar to those used in the bombings, and
that he has to be detained so that he cannot flee before the grand jury has
a chance to obtain his testimony.
May 24, 2004: The government announces that the FBI has erred in
its identification of Mayfield and moves to dismiss the material witness
proceeding.75
In March 2006, the Office of the Inspector General of the U.S. Depart-
ment of Justice issued a comprehensive analysis of how the misidentification
occurred.76 And in November 2006, the federal government agreed to pay
Mayfield $2 million for his wrongful jailing in connection with the 2004
terrorist bombings in Madrid.77 The Mayfield case and the resulting report
from the Inspector General surely signal caution against simple, and unveri-
fied, assumptions about the reliability of fingerprint evidence.
In Maryland v. Rose, a Maryland State trial court judge found that the
Analysis, Comparison, Evaluation, and Verification (ACE-V) process (see
Chapter 5) of latent print identification does not rest on a reliable factual
foundation.78 The opinion went into considerable detail about the lack of
error rates, lack of research, and potential for bias. The judge ruled that
the State could not offer testimony that any latent fingerprint matched the
prints of the defendant. The judge also noted that, because the case involved
75 S.T. Wax and C.J. Schatz. 2004. A multitude of errors: The Brandon Mayfield case. The
Champion. September-October, p. 6. The facts of the case and Mayfield’s legal claims against
the government are fully reported in Mayfield v. United States, 504 F. Supp. 2d 1023 (D. Or.
2007).
76 Office of the Inspector General, Oversight and Review Division, U.S. Department of Jus-
usdoj.gov/oig/special/s0601/exec.pdf.
77 E. Lichtblau. 2006. “U.S. Will Pay $2 Million To Lawyer Wrongly Jailed.” New York
Times. November 30, at A18.
78 Maryland v. Rose, Case No. K06-0545, mem. op. at 31 (Balt. County Cir. Ct. Oct.
19, 2007) (holding that the ACE-V methodology of latent fingerprint identification was “a
subjective, untested, unverifiable identification procedure that purports to be infallible” and
therefore ruling that fingerprint evidence was inadmissible). The ACE-V process is described
in Chapter 5.
106
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
the possibility of the death penalty, the reliability of the evidence offered
against the defendant was critically important.79
The same concerns cited by the judge in Maryland v. Rose can be raised
with respect to other forensic techniques that lack scientific validation and
careful reliability testing.
Judicial Dispositions of Questions Relating to Other Forensic Disciplines
Review of reported judicial opinions reveals that, at least in criminal
cases, forensic science evidence is not routinely scrutinized pursuant to
the standard of reliability enunciated in Daubert. The Supreme Court in
Daubert indicated that the subject of an expert’s testimony should be “sci-
entific knowledge”—which implies that such knowledge is based on sci-
entific methods—to ensure that “evidentiary reliability will be based upon
scientific validity.” The standard is admittedly “flexible,” but that does not
render it meaningless. Any reasonable reading of Daubert strongly suggests
that, when faced with forensic evidence, “trial judge[s] must ensure that any
and all scientific testimony or evidence admitted is not only relevant, but
reliable.” As the reported cases suggest, however, Daubert has done little to
improve the use of forensic science evidence in criminal cases.
For years in the forensic science community, the dominant argument
against regulating experts was that every time a forensic scientist steps
into a courtroom, his work is vigorously peer reviewed and scrutinized by
opposing counsel. A forensic scientist might occasionally make an error
in the crime laboratory, but the crucible of courtroom cross-examination
79 Professor Jennifer Mnookin has also highlighted an important concern over “the rhe-
torical dimensions of the testimony . . . provide[d] in court” by members of the fingerprint
community:
At present, fingerprint examiners typically testify in the language of absolute certainty. Both
the conceptual foundations and the professional norms of latent fingerprinting prohibit experts
from testifying to identification unless they believe themselves certain that they have made a
correct match. Experts therefore make only what they term “positive” or “absolute” identifica-
tions—essentially making the claim that they have matched the latent print to the one and only
person in the entire world whose fingertip could have produced it. In fact, if a fingerprint exam-
iner testifies on her own initiative that a match is merely “likely” or “possible” or “credible,”
rather than certain, she could possibly be subject to disciplinary sanction! Given the general lack
of validity testing for fingerprinting; the relative dearth of difficult proficiency tests; the lack of
a statistically valid model of fingerprinting; and the lack of validated standards for declaring a
match, such claims of absolute, certain confidence in identification are unjustified, the product
of hubris more than established knowledge. Therefore, in order to pass scrutiny under Daubert,
fingerprint identification experts should exhibit a greater degree of epistemological humility.
Claims of “absolute” and “positive” identification should be replaced by more modest claims
about the meaning and significance of a “match.”
J.L. Mnookin. 2008. The validity of latent fingerprint identification: Confessions of a finger-
printing moderate. Law, Probability and Risk 7(2):127; see also Koehler, supra note 71.
FORENSIC SCIENCE EVIDENCE IN LITIGATION
107
would expose it at trial. This “crucible,” however, turned out to be utterly
ineffective.
Unlike the extremely well-litigated civil challenges, the criminal defendant’s
challenge is usually perfunctory. Even when the most vulnerable forensic
sciences—hair microscopy, bite marks, and handwriting—are attacked,
the courts routinely affirm admissibility citing earlier decisions rather than
facts established at a hearing. Defense lawyers generally fail to build a
challenge with appropriate witnesses and new data. Thus, even if inclined
to mount a Daubert challenge, they lack the requisite knowledge and
skills, as well as the funds, to succeed.80
The reported decisions dealing with judicial dispositions of Daubert-
type questions appear to confirm this assessment. As noted above, the
courts often “affirm admissibility citing earlier decisions rather than facts
established at a hearing.” Much forensic evidence—including, for example,
bite marks81 and firearm and toolmark identifications82—is introduced in
80 Neufeld, supra note 44, at S109, S110.
81 There is nothing to indicate that courts review bite mark evidence pursuant to Daubert’s
standard of reliability. See, e.g., Milone v. Camp, 22 F.3d 693, 702 (7th Cir. 1994) (denying
habeas petition after finding, in part, that the inclusion of bite mark testimony against the
defendant had not denied him a fair trial, and stating that “while the science of forensic odon-
tology might have been in its infancy at the time of trial . . . certainly there is some probative
value to comparing an accused’s dentition to bite marks found on the victim.”). Two recent
cases might, at first glance, seem to indicate that courts were beginning to seriously evaluate
the general credibility of bite mark testimony, but this is not in fact the case. In Burke v. Town
of Walpole, 405 F.3d 66 (1st Cir. 2005), the court denied summary judgment to police officers
in a 42 U.S.C. § 1983 action where exculpatory DNA evidence that directly contradicted
inculpatory bite mark evidence was “intentionally or recklessly withheld from the officer who
was actually preparing the warrant application,” ibid., p. 84, resulting in petitioner being
wrongfully imprisoned for 41 days. However, the Burke court rejected the petitioner’s claim
that the inclusion of bite mark evidence in the arrest warrant had demonstrated “reckless dis-
regard for the truth,” because the method was generally unreliable. Ibid., pp. 82-83. In Ege v.
Yukins, 380 F. Supp. 2d 852 (E.D. Mich. 2005), aff’d in part and rev’d in part, 485 F.3d 364
(6th Cir. 2007), the court granted the habeas petition of a defendant whose conviction was
based in significant part on bite mark testimony from a later-discredited expert witness. But the
disposition in Ege rested primarily on the flaws of one “particular witness and his particular
testimony,” not on a judicial evaluation of “the [bite mark] field’s more general shortcomings.”
4 Faigman et al., op. cit., supra note 1, § 36:6, p. 662.
82 There is little to indicate that courts review firearms evidence pursuant to Daubert’s stan-
dard of reliability. See e.g., United States v. Hicks, 389 F.3d 514 (5th Cir. 2004) (upholding
defendant’s conviction after finding, in part, that it was not an abuse of discretion for the court
to admit testimony on shell casing comparisons by the Government’s firearms expert); United
States v. Foster, 300 F. Supp. 2d 375 (D. Md. 2004) (denying defendant’s motion to exclude
expert firearms testimony). Several federal trial judges, however, have subjected expert firearm
testimony to rigorous analysis under Daubert. In United States v. Monteiro, 407 F. Supp.
2d 351 (D. Mass. 2006), Judge Saris concluded that toolmark identification testimony was
108
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
criminal trials without any meaningful scientific validation, determination
of error rates, or reliability testing to explain the limits of the discipline.
One recent judicial decision highlights the problem. In United States v.
Green, Judge Gertner acknowledged that toolmark identification testi-
mony ought not be considered admissible under Daubert.83 But the judge
pointed out that “the problem for the defense is that every single court
post-Daubert has admitted this testimony, sometimes without any search-
ing review, much less a hearing.”84 Judge Gertner allowed the prosecution’s
expert to describe the similarities between the shell casings at issue, but
prohibited him from testifying that there was a definitive match. Obviously
feeling bound by circuit precedent, the judge stated:
I reluctantly [admit the evidence] because of my confidence that any other
decision will be rejected by appellate courts, in light of precedents across
the country, regardless of the findings I have made. While I recognize
that the Daubert-Kumho standard does not require the illusory perfec-
tion of a television show (CSI, this wasn’t), when liberty hangs in the
balance—and, in the case of the defendants facing the death penalty, life
itself—the standards should be higher than were met in this case, and than
have been imposed across the country. The more courts admit this type of
toolmark evidence without requiring documentation, proficiency testing,
or evidence of reliability, the more sloppy practices will endure; we should
require more.85
“[T]he undeniable reality is that the community of forensic science
generally admissible under Daubert, but excluded the specific testimony at issue, because the
experts failed to properly document their basis for identification, and because an independent
examiner had not verified the experts’ conclusions. Likewise, in United States v. Diaz, No.
05-CR-167, 2007 WL 485967, at *14 (N.D. Cal. Feb. 12, 2007), Judge Alsup allowed firearm
identification testimony under Daubert, but prevented experts from testifying to their conclu-
sions “to the exclusion of all other firearms in the world” and only allowed testimony “to a
reasonable degree of certainty.” Cf. United States v. Glynn, 578 F. Supp. 2d 569 (S.D.N.Y.
2008), where Judge Rakoff precluded testimony that a bullet and shell casings came from
a firearm linked to the defendant “to a reasonable degree of ballistics certainty,” because
“whatever else ballistics identification analysis could be called, it could not fairly be called
‘science.’” However, the judge ruled that although inadmissible under Daubert, testimony that
the evidence was “more likely than not” from the firearm was admissible under Federal Rule
of Evidence 401. See also Green, 405 F. Supp. 2d 104, discussed in the text.
83 405 F. Supp. 2d at 107-08.
84 Ibid., p. 108.
85 Ibid., p. 109 (footnotes omitted). “The case law on the admissibility of toolmark iden-
tification and firearms identification expert evidence is typified by decisions admitting such
testimony with little, and usually no, reference to legal authority beyond broad ‘discretion’ and
an adroit sidestepping of any judicial duty to assure that experts’ claims are valid. Appellate
courts defer to trial courts, and trial courts defer to juries. Later appellate courts simply defer
to earlier appellate courts.” 4 Faigman et al., op. cit., supra note 1, § 34:5, p. 589.
FORENSIC SCIENCE EVIDENCE IN LITIGATION
109
professionals has not done nearly as much as it reasonably could have
done to establish either the validity of its approach or the accuracy of its
practitioners’ conclusions,”86 and the courts have been “utterly ineffective”
in addressing this problem.87
CONCLUSION
Prophetically, the Daubert decision observed that “there are important
differences between the quest for truth in the courtroom and the quest for
truth in the laboratory. Scientific conclusions are subject to perpetual revi-
sion. Law, on the other hand, must resolve disputes finally and quickly.”88
But because accused parties in criminal cases are convicted on the basis of
testimony from forensic science experts, much depends upon whether the
evidence offered is reliable. Furthermore, in addition to protecting innocent
persons from being convicted of crimes that they did not commit, we are
also seeking to protect society from persons who have committed criminal
acts. Law enforcement officials and the members of society they serve need
to be assured that forensic techniques are reliable. Therefore, we must limit
the risk of having the reliability of certain forensic science methodologies
condoned by the courts before the techniques have been properly studied
and their accuracy verified. “[T]here is no evident reason why [‘rigorous,
systematic’] research would be infeasible.”89 However, some courts appear
to be loath to insist on such research as a condition of admitting forensic
science evidence in criminal cases, perhaps because to do so would likely
“demand more by way of validation than the disciplines can presently
offer.”90
Some legal scholars think that, “[o]ver time, if Daubert does not come
86 Mnookin, op. cit., supra note 79.
87 Neufeld, op. cit., supra note 44, p. S109. In Green, 405 F. Supp. 2d at 109 n.6, Judge
Gertner also noted that:
[R]ecent reexaminations of relatively established forensic testimony have produced striking
results. Saks and Koehler, for example, report that forensic testing errors were responsible for
wrongful convictions in 63% of the 86 DNA Exoneration cases reported by the Innocence Proj-
ect at Cardozo Law School. Michael Saks and Jonathan Koehler, The Coming Paradigm Shift
in Forensic Identification Science, 309 Science 892 (2005). This only reinforces the importance
of careful analysis of expert testimony in this case.
See also S.R. Gross, Convicting the Innocent (U. Mich. Law Sch. Pub. Law & Legal Theory
Working Paper Series, Working Paper No. 103, 2008). Available at http://papers.ssrn.com/
sol3/papers.cfm?abstract_id=1100011 (forthcoming in Annual Review of Law & Social Sci-
ence 2008).
88 Daubert v. Merrell Dow Pharm., Inc., 509 U.S. 579, 596-97 (1993).
89 J. Griffin and D.J. LaMagna. 2002. Daubert challenges to forensic evidence: Ballistics
next on the firing line. The Champion. September-October:21.
90 Ibid. See, e.g., Crisp, 324 F.3d at 270.
110
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
to be diluted or distorted, . . . courts will increasingly appreciate its power
and flexibility to evaluate proffered expert testimony.”91 However, at least
with respect to criminal cases, this may reflect an unrealistic assessment of
the problem. “The principal difficulty, it appears, is that many [forensic
science] techniques have been relied on for so long that courts might be re-
luctant to rethink their role in the trial process
In many forensic areas,
effectively no research exists to support the practice.”92
As the discussion in this chapter indicates, the adversarial process re-
lating to the admission and exclusion of scientific evidence is not suited to
the task of finding “scientific truth.” The judicial system is encumbered by,
among other things, judges and lawyers who generally lack the scientific
expertise necessary to comprehend and evaluate forensic evidence in an
informed manner, trial judges (sitting alone) who must decide evidentiary
issues without the benefit of judicial colleagues and often with little time
for extensive research and reflection, and the highly deferential nature of
the appellate review afforded trial courts’ Daubert rulings. Furthermore,
the judicial system embodies a case-by-case adjudicatory approach that is
not well suited to address the systematic problems in many of the various
forensic science disciplines. Given these realities, there is a tremendous
need for the forensic science community to improve. Judicial review, by
itself, will not cure the infirmities of the forensic science community.93 The
development of scientific research, training, technology, and databases asso-
ciated with DNA analysis have resulted from substantial and steady federal
support for both academic research and programs employing techniques
for DNA analysis. Similar support must be given to all credible forensic
science disciplines if they are to achieve the degrees of reliability needed
to serve the goals of justice. With more and better educational programs,
accredited laboratories, certified forensic practitioners, sound operational
principles and procedures, and serious research to establish the limits and
measures of performance in each discipline, forensic science experts will be
better able to analyze evidence and coherently report their findings in the
courts. The present situation, however, is seriously wanting, both because
of the limitations of the judicial system and because of the many problems
faced by the forensic science community.
91 1 Faigman et al., op. cit., supra note 1, § 1:1, p. 5 n. 9.
92 Ibid. § 1:30, p. 85 (footnotes omitted).
93 See J.L. Mnookin. Expert evidence, partisanship, and epistemic competence. 73 Brook.
L. Rev. 1009, 1033 (2008) (“[S]o long as we have our adversarial system in much its pres-
ent form, we are inevitably going to be stuck with approaches to expert evidence that are
imperfect, conceptually unsatisfying, and awkward. It may well be that the real lesson is this:
those who believe that we might ever fully resolve—rather than imperfectly manage—the
deep structural tensions surrounding both partisanship and epistemic competence that per-
meate the use of scientific evidence within our legal system are almost certainly destined for
disappointment.”).
4
The Principles of Science and
Interpreting Scientific Data
Scientific method refers to the body of techniques for investigating phe-
nomena, acquiring new knowledge, or correcting and integrating previous
knowledge. It is based on gathering observable, empirical and measurable
evidence subject to specific principles of reasoning.
Isaac Newton (1687, 1713, 1726)
“Rules for the study of natural philosophy,”
Philosophiae Naturalis Principia Mathematica
Forensic science actually is a broad array of disciplines, as will be
seen in the next chapter. Each has its own methods and practices, as well
as its strengths and weaknesses. In particular, each varies in its level of
scientific development and in the degree to which it follows the principles
of scientific investigation. Adherence to scientific principles is important
for concrete reasons: they enable the reliable inference of knowledge from
uncertain information—exactly the challenge faced by forensic scientists.
Thus, the reliability of forensic science methods is greatly enhanced when
those principles are followed. As Chapter 3 observes, the law’s admission
of and reliance on forensic evidence in criminal trials depends critically on
(1) the extent to which a forensic science discipline is founded on a reliable
scientific methodology, leading to accurate analyses of evidence and proper
reports of findings and (2) the extent to which practitioners in those foren-
sic science disciplines that rely on human interpretation adopt procedures
and performance standards that guard against bias and error. This chapter
discusses the ways in which science more generally addresses those goals.
111
112
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
FUNDAMENTAL PRINCIPLES OF THE SCIENTIFIC METHOD
The scientific method presumes that events occur in consistent patterns
that can be understood through careful comparison and systematic study.
Knowledge is produced through a series of steps during which data are
accumulated methodically, strengths and weaknesses of information are as-
sessed, and knowledge about causal relationships is inferred. In the process,
scientists also develop an understanding of the limits of that knowledge
(such as the precision of the observations), the inferred nature of relation-
ships, and key assumptions behind the inferences. Hypotheses are devel-
oped, are measured against the data, and are either supported or refuted.
Scientists continually observe, test, and modify the body of knowledge.
Rather than claiming absolute truth, science approaches truth either through
breakthrough discoveries or incrementally, by testing theories repeatedly.
Evidence is obtained through observations and measurements conducted
in the natural setting or in the laboratory. In the laboratory, scientists can
control and vary the conditions in order to isolate exclusive effects and
thus better understand the factors that influence certain outcomes. Typi-
cally, experiments or observations must be conducted over a broad range of
conditions before the roles of specific factors, patterns, or variables can be
understood. Methods to reduce errors are part of the study design, so that,
for example, the size of the study is chosen to provide sufficient statistical
power to draw conclusions with a high level of confidence or to understand
factors that might confound results. Throughout scientific investigations,
the investigator must be as free from bias as possible, and practices are put
in place to detect biases (such as those from measurements, human inter-
pretation) and to minimize their effects on conclusions.
Ultimately, the goal is to construct explanations (“theories”) of phe-
nomena that are consistent with broad scientific principles, such as the
laws of thermodynamics or of natural selection. These theories, and in-
vestigations of them through experiments and observed data, are shared
through conferences, publications, and collegial interactions, which push
the scientist to explain his or her work clearly and which raise questions
that might not have been considered. The process of sharing data and re-
sults requires careful recordkeeping, reviewed by others. In addition, the
need for credibility among peers drives investigators to avoid conflicts of
interest. Acceptance of the work comes as results and theories continue to
hold, even under the scrutiny of peers, in an environment that encourages
healthy skepticism. That scrutiny might extend to independent reproduc-
tion of the results or experiments designed to test the theory under different
conditions. As credibility accrues to data and theories, they become ac-
cepted as established fact and become the “scaffolding” upon which other
investigations are constructed.
THE PRINCIPLES OF SCIENCE
113
This description of how science creates new theories illustrates key ele-
ments of good scientific practice: precision when defining terms, processes,
context, results, and limitations; openness to new ideas, including criticism
and refutation; and protections against bias and overstatement (going be-
yond the facts). Although these elements have been discussed here in the
context of creating new methods and knowledge, the same principles hold
when applying known processes or knowledge. In day-to-day forensic sci-
ence work, the process of formulating and testing hypotheses is replaced
with the careful preparation and analysis of samples and the interpretation
of results. But that applied work, if done well, still exhibits the same hall-
marks of basic science: the use of validated methods and care in following
their protocols; the development of careful and adequate documentation;
the avoidance of biases; and interpretation conducted within the constraints
of what the science will allow.
Validation of New Methods
One particular task of science is the validation of new methods to
determine their reliability under different conditions and their limitations.
Such studies begin with a clear hypothesis (e.g., “new method X can
reliably associate biological evidence with its source”). An unbiased ex-
periment is designed to provide useful data about the hypothesis. Those
data—measurements collected through methodical prescribed observations
under well-specified and controlled conditions—are then analyzed to sup-
port or refute the hypothesis. The thresholds for supporting or refuting the
hypothesis are clearly articulated before the experiment is run. The most
important outcomes from such a validation study are (1) information about
whether or not the method can discriminate the hypothesis from an alter-
native, and (2) assessments of the sources of errors and their consequences
on the decisions returned by the method. These two outcomes combine to
provide precision and clarity about what is meant by “reliably associate.”
For a method that has not been subjected to previous extensive study, a
researcher might design a broad experiment to assist in gaining knowledge
about its performance under a range of conditions. Those data are then
analyzed for any underlying patterns that may be useful in planning or
interpreting tests that use the new method. In other situations, a process
already has been formulated from existing experimental data, knowledge,
and theory (e.g., “biological markers A, B, and C can be used in DNA
forensic investigations to pair evidence with suspect”).
To confirm the validity of a method or process for a particular purpose
(e.g., for a forensic investigation), validation studies must be performed.
The International Organization for Standardization (ISO) and the In-
ternational Electrotechnical Commission (IEC) developed a joint document,
114
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
“General requirements for the competence of testing and calibration labo-
ratories” (commonly referred to as “ISO 17025”), which includes a well-
established list of techniques that can be used, alone or in combination, to
validate a method:
• calibration using reference standards or reference materials;
• comparison of results achieved with other methods;
• interlaboratory comparisons;
• systematic assessment of the factors influencing the result; and
• assessment of the uncertainty of the results based on scientific un-
derstanding of the theoretical principles of the method and practi-
cal experience.1
A critical step in such validation studies is their publication in peer-
reviewed journals, so that experts in the field can review, question, and
check the repeatability of the results. These publications must include clear
statements of the hypotheses under study, as well as sufficient details about
the experiments, the resulting data, and the data analysis so that the studies
can be replicated. Replication will expose not only additional sources of
variability but also further aspects of the process, leading to greater under-
standing and scientific knowledge that can be used to improve the method.
Methods that are specified in more detail (such as DNA analysis, where
particular genetic loci are to be compared) will have greater credibility and
also are more amenable to systematic improvement than those that rely
more heavily on the judgments of the investigator.
The validation of results over time increases confidence. Moreover,
the scientific culture encourages continued questioning and improvement.
Thus, the relevant scientific community continues to check that established
results still hold under new conditions and that they continue to hold in the
face of new knowledge. The involvement of graduate student researchers in
scientific research contributes greatly to this diligence, because part of their
education is to read carefully and to question so-called established methods.
This culture leads to continued reexamination of past research and hence
increased knowledge.
In the case of DNA analysis, studies have evaluated the precision, reli-
ability, and uncertainties of the methods. This knowledge has been used to
define standard procedures that, when followed, lead to reliable evidence.
For example, below is a brief sample of the specifications required by the
Federal Bureau of Investigation’s (FBI’s) Quality Assurance Standards for
1 Quoted from Section 5.4.5 2 (Note 2) of ISO/IEC 17025, “General requirements for the
competence of testing and calibration laboratories” (2nd ed., May 15, 2005).
THE PRINCIPLES OF SCIENCE
115
Forensic DNA Testing Laboratories2 in order to ensure reliable DNA fo-
rensic analysis:
•
Testing laboratories must have a standard operating protocol for
each analytical technique used, specifying reagents, sample prepa-
ration, extraction, equipment, and controls that are standard for
DNA analysis and data interpretation.
•
The laboratory shall monitor the analytical procedures using ap-
propriate controls and standards, including quantitation standards
that estimate the amount of human nuclear DNA recovered by ex-
traction, positive and negative amplification controls, and reagent
blanks.
•
The laboratory shall check its DNA procedures annually or when-
ever substantial changes are made to the protocol(s) against an
appropriate and available NIST standard reference material or
standard traceable to a NIST standard.
•
The laboratory shall have and follow written general guidelines for
the interpretation of data.
•
The laboratory shall verify that all control results are within estab-
lished tolerance limits.
•
Where appropriate, visual matches shall be supported by a numeri-
cal match criterion.
•
For a given population(s) and/or hypothesis of relatedness, the
statistical interpretation shall be made following the recommenda-
tions 4.1, 4.2, or 4.3 as deemed applicable of the National Research
Council report entitled The Evaluation of Forensic DNA Evidence
(1996) and/or a court-directed method. These calculations shall be
derived from a documented population database appropriate for
the calculation.3
This level of specificity is consistent with the spirit of the guidelines
presented in ISO 17025. The second edition (May 15, 2005) of those
guidelines includes the following minimum set of information for properly
specifying the process of any new analytical method:
(a) appropriate identification;
(b) scope;
(c) description of the type of item to be tested or calibrated;
bioforensics.com/conference04/TWGDAM/Quality_Assurance_Standards_2.pdf.
3 Paraphrased from Section 9 of the FBI’s Quality Assurance Standards for Forensic DNA
Testing Laboratories.
116
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
(d)
parameters or quantities and ranges to be determined;
(e)
apparatus and equipment, including technical performance
requirements;
(f)
reference standards and reference materials required;
(g)
environmental conditions required and any stabilization period
needed;
(h)
description of the procedure, including
-
affixing of identification marks, handling, transporting, storing
and preparation of items;
-
checks to be made before the work is started;
-
checks that the equipment is working properly and, where
required, calibration and adjustment of the equipment before
each use;
-
the method of recording the observations and results;
-
any safety measures to be observed;
(i)
criteria and/or requirements for approval/rejection;
(j)
data to be recorded and method of analysis and presentation;
(k)
the uncertainty or the procedure for estimating uncertainty.4
Uncertainty and Error
Scientific data and processes are subject to a variety of sources of error.
For example, laboratory results and data from questionnaires are subject to
measurement error, and interpretations of evidence by human observers are
subject to potential biases. A key task for the scientific investigator design-
ing and conducting a scientific study, as well as for the analyst applying a
scientific method to conduct a particular analysis, is to identify as many
sources of error as possible, to control or to eliminate as many as possible,
and to estimate the magnitude of remaining errors so that the conclusions
drawn from the study are valid. Numerical data reported in a scientific
paper include not just a single value (point estimate) but also a range of
plausible values (e.g., a confidence interval, or interval of uncertainty).
Measurement Error
As with all other scientific investigations, laboratory analyses con-
ducted by forensic scientists are subject to measurement error. Such error
reflects the intrinsic strengths and limitations of the particular scientific
technique. For example, methods for measuring the level of blood alcohol
in an individual or methods for measuring the heroin content of a sample
4 Quoted from Section 5.4.4 of ISO/IEC 17025, “General requirements for the competence
of testing and calibration laboratories” (2nd ed., May 15, 2005).
THE PRINCIPLES OF SCIENCE
117
can do so only within a confidence interval of possible values. In addi-
tion to the inherent limitations of the measurement technique, a range of
other factors may also be present and can affect the accuracy of laboratory
analyses. Such factors may include deficiencies in the reference materials
used in the analysis, equipment errors, environmental conditions that lie
outside the range within which the method was validated, sample mix-ups
and contamination, transcriptional errors, and more.
Consider, for example, a case in which an instrument (e.g., a breatha-
lyzer such as Intoxilyzer) is used to measure the blood-alcohol level of an
individual three times, and the three measurements are 0.08 percent, 0.09
percent, and 0.10 percent. The variability in the three measurements may
arise from the internal components of the instrument, the different times
and ways in which the measurements were taken, or a variety of other fac-
tors. These measured results need to be reported, along with a confidence
interval that has a high probability of containing the true blood-alcohol
level (e.g., the mean plus or minus two standard deviations). For this il-
lustration, the average is 0.09 percent and the standard deviation is 0.01
percent; therefore, a two-standard-deviation confidence interval (0.07 per-
cent, 0.11 percent) has a high probability of containing the person’s true
blood-alcohol level. (Statistical models dictate the methods for generating
such intervals in other circumstances so that they have a high probability of
containing the true result.) The situation for assessing heroin content from
a sample of white powder is similar, although the quantification and limits
are not as broadly standardized. The combination of gas chromatography
and mass spectrometry (GC/MS) is used extensively in identifying con-
trolled substances. Those analyses tend to be more qualitative (e.g., iden-
tifying peaks on a spectrum that appear at frequencies consistent with the
controlled substance and which stand out above the background “noise”),
although quantification is possible.
Error Rates
Analyses in the forensic science disciplines are conducted to provide
information for a variety of purposes in the criminal justice process. How-
ever, most of these analyses aim to address two broad types of questions:
(1) can a particular piece of evidence be associated with a particular class
of sources? and (2) Can a particular piece of evidence be associated with
one particular source? The first type of question leads to “classification”
conclusions. An example of such a question would be whether a particular
hair specimen shares physical characteristics common to a particular ethnic
group. An affirmative answer to a classification question indicates only that
the item belongs to a particular class of similar items. Another example
might be whether a paint mark left at a crime scene is consistent (according
118
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
to some collection of relevant measurements) with a particular paint sample
in a database, from which one can infer the class of vehicle (e.g., model(s)
and production year(s)) that could have left the mark. The second type of
question leads to “individualization” conclusions—for example, does a
particular DNA sample belong to individual X?
Although the questions addressed by forensic analyses are not always
binary (yes/no) or as crisply stated as in the previous paragraph, the para-
digm of yes/no conclusions is useful for describing and quantifying the
accuracy with which forensic science disciplines can provide answers.5 In
such situations, results from analyses for which the truth is known can be
classified in a two-way table as follows:
Analysis Results
Truth
yes
no
yes
a (true positives)
b (false negatives)
no
c (false positives)
d (true negatives)
The conceptual framework and terminology for evaluating the accu-
racy of forensic analyses is illustrated using a hypothetical example from
microscopic analysis of head hair. In this situation, multiple features, both
qualitative and quantitative, on each sample of hair are assessed. Qualita-
tive features include color (e.g., blonde, brown, red), coloring (natural or
treated), form (straight, wavy, curved, kinked), texture (smooth, medium,
coarse). Quantitative features include length and diameter. Undoubtedly,
these features will vary from hair to hair, even from the same individual,
but features that vary less for the same individual (i.e., within-individual
variability) and more for different individuals (i.e., between-individual vari-
ability) are needed for purposes of class identification and discrimination.
These features may also be combined in some fashion to result in some
overall score, or set of scores, for each sample, and these scores are then
compared with those from the target sample. In the final analysis, however,
a binary conclusion is often required. For example, “Did this hair come
from the head of a Caucasian person?”
As in the case of all analyses leading to classification conclusions (e.g.,
diagnostic tests in medicine), the microscopic hair analysis process must
be subjected to performance and validation studies in which appropriate
error rates can be defined and estimated. Consider a hypothetical study in
5 More complete discussion of the questions addressed by forensic science may be found
in references such as K. Inman and N. Rudin. 2002. The origin of evidence. Forensic Science
International 126:11-16; and R. Cook, I.W. Evett, G. Jackson, P.J. Jones, and J.A. Lambert.
1998. A hierarchy of propositions: Deciding which level to address in casework. Science and
Justice 38:231-239.
THE PRINCIPLES OF SCIENCE
119
which 100 samples (each with multiple hairs) are taken from the heads of
100 individuals from class C, and another 100 samples are taken from the
heads of individuals not in class C. The analyst is asked to determine, for
each of the 200 samples, whether it does or does not come from a person
in class C, and the true answer is known. The validation study returns the
following results:
Hypothetical Hair Analysis Validation Study
Analysis of Hair Samples Indicates:
Class C
Not Class C
Row Total
Sample is from Class
95
5
100
C Persons
True Positive (correct
False Negative
determination)
Sample is not from
2
98
100
Class C Persons
False Positive
True Negative
(correct
determination)
Column Total
97
103
Overall total
200
The accuracy of a test (here, microscopic hair analysis) can be assessed
in different ways. Borrowing terminology from the evaluation of medical
diagnostic tests, four characterizations and their associated measures are
given below. Each one is useful in its own way: the first two emphasize the
ability to detect an association; the last two emphasize the ability to predict
an association:6
• Among samples from persons in Class C, the fraction that is cor-
rectly identified by the test is called the “sensitivity” or the “true
positive rate” (TPR) of the test. In this table, the sensitivity would
be estimated as [95/(95+5)] × 100=95 percent.
• Among samples from persons not in Class C, the fraction that is
correctly identified by the test is called the “specificity” or the “true
6 See, e.g., X-H. Zhou, N. Obuchowski, and D. McClish. 2002. Statistical Methods in
Diagnostic Medicine. Hoboken, NJ. Wiley & Sons, for a general account of methods for
diagnostic tests. A series of NAS/NRC reports have applied such methods to the examination
of forensic disciplines. See, e.g., NRC. Committee to Review the Scientific Evidence on the
Polygraph. 2003. The Polygraph and Lie Detection. Washington, DC: The National Acad-
emies Press; NRC. 2004. Forensic Analysis: Weighing Bullet Lead Evidence. Washington, DC:
The National Academies Press; NAS. 2005. The Sackler Colloquium on Forensic Science: The
Nexus of Science and the Law, November 16-18, 2005.
120
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
negative rate” (TNR) of the test. In this table, the specificity would
be estimated as [98/(2+98)] × 100=98 percent.
• Among samples classified by the test as coming from persons in
Class C, the fraction that actually turns out to be from Class C is
called the “positive predictive value (PPV)” of the test. In this table,
the PPV would be estimated as [95/(95+ 2)] × 100=98 percent.
• Among samples classified by the test as coming from persons not in
Class C, the fraction that actually turns out to not be persons from
Class C is called the “negative predictive value (NPV)” of the test.
In this table, the NPV would be estimated as [98/(5+98)] × 100=95
percent.
The above four measures emphasize the ability of the analysis to make
correct determinations.7 “Error rates” are defined as proportions of cases in
which the analysis led to a false conclusion. For example, the complement
of sensitivity (100 percent minus the sensitivity) is the percent of false nega-
tive cases in which the sample was from class C but the analysis reached
the opposite conclusion. In the above table, this would be estimated as 5
percent. Similarly, the complement of specificity (100 percent minus the
specificity) is the percent of false positive cases in which the sample was
not from class C but the analysis concluded that it was. In the above table
this would be estimated as 2 percent. A global error rate could be defined
as the percent of incorrectly identified cases among all those analyzed. In
the above table this would be estimated as [(5+2)/200] × 100=3.5 percent.
Importantly, whether the test answer is correct or not depends on which
question is being addressed by the test. In this hair comparison example,
the purpose is to determine whether the hair came from the head of an
individual from class C. Thus, the analysis should be evaluated on the ac-
curacy of the classification. In this example, if the analysis indicated “Class
C” but the hair actually came from a “non-Class C” individual, then the
analysis returned an incorrect classification. This accuracy evaluation does
not apply to other tasks that are beyond the goal of the particular analysis,
such as pinpointing the individual from whom the specimen was obtained.
In the paint example about paint marks left by a vehicle, if the question is
whether a vehicle under investigation was a model A made by manufacturer
B in 2000, then a correct answer is limited to only the model, manufacturer,
and year.
7 Each estimate (of sensitivity, specificity, PPV, NPV) is associated with an interval that
has a high probability of containing the true sensitivity, specificity, PPV, NPV. The larger the
study, the more precise the estimate (i.e., the narrower the interval of uncertainty about the
estimate).
THE PRINCIPLES OF SCIENCE
121
Although only illustrations, these examples serve to demonstrate the
importance of:
• the careful and precise characterization of the scientific procedure,
so that others can replicate and validate it;
• the identification of as many sources of error as possible that can
affect both the accuracy and precision of a measurement;
• the quantification of measurements
(e.g., in the example of
GC/MS analysis of possible heroin, reporting peak area, as well
as appropriate calibration data, including the response area for a
known amount of analyte standard, rather than merely “peak is
present/absent”);
• the reporting of a measurement with an interval that has a high
probability of containing the true value;
• the precise definition of the question addressed by the method (e.g.,
classification versus individualization), and the recognition of its
limitations; and
• the conducting of validation studies of the performance of a foren-
sic procedure to assess the percentages of false positives and false
negatives.
Clearly, better understanding of the measuring equipment and the
measurement process leads to more improvements to every process and
ultimately to fewer false positive and false negative results. Most impor-
tantly, as stated above, whether the test answer is correct or not depends
on the question the test is being used to address. In the case of microscopic
hair analysis, the validation study may confirm its value in identifying class
characteristics of an individual, but not in identifying the specific person.
It is also important to note that errors and corresponding error rates
can have more complex sources than can be accommodated within the
simple framework presented above. For example, in the case of DNA
analysis, a declaration that two samples match can be erroneous in at least
two ways: The two samples might actually come from different individuals
whose DNA appears to be the same within the discriminatory capability of
the tests, or two different DNA profiles could be mistakenly determined to
be matching. The probability of the former error is typically very low, while
the probability of a false positive (different profiles wrongly determined to
be matching) may be considerably higher. Both sources of error need to be
explored and quantified in order to arrive at reliable error rate estimates
for DNA analysis.8
8 C. Aitken and F. Taroni. 2004. Statistics and the Evaluation of Evidence for Forensic
Scientists. Chichester, UK: John Wiley & Sons.
122
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
The existence of several types of potential error rates makes it abso-
lutely critical for all involved in the analysis to be explicit and precise in
the particular rate or rates referenced in a specific setting. The estimation
of such error rates requires rigorously developed and conducted scientific
studies. Additional factors may play a role in analyses involving human
interpretation, such as the experience, training, and inherent ability of the
interpreter, the protocol for conducting the interpretation, and biases from
a variety of sources, as discussed in the next section. The assessment of the
accuracy of the conclusions from forensic analyses and the estimation of
relevant error rates are key components of the mission of forensic science.
Sources of Bias
Human judgment is subject to many different types of bias, because we
unconsciously pick up cues from our environment and factor them in an
unstated way into our mental analyses. Those mental analyses might also
be affected by unwarranted assumptions and a degree of overconfidence
that we do not even recognize in ourselves. Such cognitive biases are not
the result of character flaws; instead, they are common features of deci-
sionmaking, and they cannot be willed away.9 A familiar example is how
the common desire to please others (or avoid conflict) can skew one’s judg-
ment if co-workers or supervisors suggest that they are hoping for, or have
reached, a particular outcome. Science takes great pains to avoid biases by
using strict protocols to minimize their effects. The 1996 National Acad-
emies DNA report, for example, notes, “[l]aboratory procedures should be
designed with safeguards to detect bias and to identify cases of true ambigu-
ity. Potential ambiguities should be documented.”10
A somewhat obvious cognitive bias that may arise in forensic science
is a willingness to ignore base rate information in assessing the probative
value of information. For example, suppose carpet fibers from a crime scene
are found to match carpet fibers found in a suspect’s home. The probative
value of this information depends on the rate at which such fibers are found
in homes in addition to that of the suspect. If the carpet fibers are extremely
common, the presence of matching fibers in the suspect’s home will be of
little probative value.11
A common cognitive bias is the tendency for conclusions to be affected
by how a question is framed or how data are presented. In a police line-up,
9 See, e.g., M.J. Saks, D.M. Risinger, R. Rosenthal, and W.C. Thompson. 2003. Context ef-
fects in forensic science: A review and application of the science of science to crime laboratory
practice in the United States. Science and Justice 43(2):77-90.
10 NRC. 1996. The Evaluation of Forensic DNA Evidence. Washington, DC: National
Academy Press.
11 C. Guthrie, J.J. Rachlinski, and A.J. Wistrich. 2001. Inside the judicial mind. Cornell
Law Review 86:777-830.
THE PRINCIPLES OF SCIENCE
123
for instance, an eyewitness who is presented with a pool of faces in one
batch might assume that the suspect is among them, which may not be cor-
rect. If the mug shots are presented together at one time and the witness is
asked to identify the suspect, the witness may choose the photograph that
is most similar to the perpetrator, even if the perpetrator’s picture is not
among those presented. Similarly, if the photographs are presented sequen-
tially and the witness knows that only a limited number will be presented,
the eyewitness might tend to “identify” one of the last photographs under
the assumption that the suspect must be in that batch. (This is also driven
by the common bias toward reaching closure.) A series of studies has shown
that judges can be subject to errors in judgment resulting from similar cog-
nitive biases.12 Forensic scientists also can be affected by this cognitive bias
if, for example, they are asked to compare two particular hairs, shoeprints,
fingerprints—one from the crime scene and one from a suspect—rather than
comparing the crime scene exemplar with a pool of counterparts.
Another potential bias is illustrated by the erroneous fingerprint iden-
tification of Brandon Mayfield as someone involved with the Madrid train
bombing in 2004. The FBI investigation determined that once the finger-
print examiner had declared a match, both he and other examiners who
were aware of this finding were influenced by the urgency of the investiga-
tion to affirm repeatedly this erroneous decision.13
Recent research provided additional evidence of this sort of bias
through an experiment in which experienced fingerprint examiners were
asked to analyze fingerprints that, unknown to them, they had analyzed
previously in their careers. For half the examinations, contextual biasing
was introduced. For example, the instructions accompanying the latent
prints included information such as the “suspect confessed to the crime”
or the “suspect was in police custody at the time of the crime.” In 6 of the
24 examinations that included contextual manipulation, the examiners
reached conclusions that were consistent with the biasing information and
different from the results they had reached when examining the same prints
in their daily work.14
Other cognitive biases may be traced to common imperfections in our
reasoning ability. One commonly recognized bias is the tendency to avoid
cognitive dissonance, such as persuading oneself through rational argu-
ment that a purchase was a good value once the transaction is complete. A
scientist encounters this unconscious bias if he/she becomes too wedded to
a preliminary conclusion, so that it becomes difficult to accept new infor-
12 Ibid.
13 R.B. Stacey. 2004. A report on the erroneous fingerprint individualization in the Madrid
train bombing case. Journal of Forensic Identification 54:707.
14 I.E. Dror and D. Charlton. 2006. Why experts make errors. Journal of Forensic Identi-
fication 56(4):600-616.
124
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
mation fairly and unduly difficult to conclude that the initial hypotheses
were wrong. This is often manifested by what is known as “anchoring,”
the well-known tendency to rely too heavily on one piece of information
when making decisions. Often, the piece of information that is weighted
disproportionately is one of the very first ones encountered. One tends to
seek closure and to view the initial part of an investigation as a “sunk cost”
that would be wasted if overturned.
Another common cognitive bias is the tendency to see patterns that do
not actually exist. This bias is related to our tendency to underestimate the
amount of complexity that can really exist in nature. Both tendencies can
lead one to formulate overly simple models of reality and thus to read too
much significance into coincidences and surprises. More generally, human
intuition is not a good substitute for careful reasoning when probabilities
are concerned. As an example, consider a problem commonly posed in
beginning statistics classes: How many people must be in a room before
there is a 50 percent probability that at least two will share a common
birthday? Intuition might suggest a large number, perhaps over 100, but
the actual answer is 23. This is not difficult to prove through careful logic,
but intuition is likely to be misleading.
All of these sources of bias are well known in science, and a large
amount of effort has been devoted to understanding and mitigating them.
The goal is to make scientific investigations as objective as possible so the
results do not depend on the investigator. Certain fields of science (most
notably, biopharmaceutical clinical trials of treatment protocols and drugs)
have developed practices such as double-blind tests and independent (blind)
verification to minimize the impact of biases. Additionally, science seeks to
publish its discoveries, findings, and conclusions so that they are subjected
to independent peer review; this enables others to study biases that may
exist in the investigative method or attempt to replicate unexpected results.
Avoiding, or compensating for, a bias is an important task. Even fields
with well-established protocols to minimize the effects of bias can still bear
improvement. For example, a recent working paper15 has raised questions
about the way cognitive dissonance has been studied since 1956. Although
these results must be considered preliminary because the paper has yet to
be published, they do demonstrate that continual vigilance is needed. Re-
search has been sparse on the important topic of cognitive bias in forensic
science—both regarding their effects and methods for minimizing them.16
15 M.K. Chen. 2008. Rationalization and Cognitive Dissonance: Do Choices Affect or
Reflect Preferences? Available at www.som.yale.edu/Faculty/keith.chen/papers/CogDisPaper.
pdf.
16 See, e.g., I.E. Dror, D. Charlton, and A.E. Peron. 2006. Contextual information renders
experts vulnerable to making erroneous identifications. Forensic Science International 156:74-
78; I.E. Dror, A. Peron, S. Hind, and D. Charlton. 2005. When emotions get the better of us:
THE PRINCIPLES OF SCIENCE
125
The Self-Correcting Nature of Science
The methods and culture of scientific research enable it to be a self-
correcting enterprise. Because researchers are, by definition, creating new
understanding, they must be as cautious as possible before asserting a new
“truth.” Also, because researchers are working at a frontier, few others
may have the knowledge to catch and correct any errors they make. Thus,
science has had to develop means of revisiting provisional results and re-
vealing errors before they are widely used. The processes of peer review,
publication, collegial interactions (e.g., sharing at conferences), and the in-
volvement of graduate students (who are expected to question as they learn)
all support this need. Science is characterized also by a culture that encour-
ages and rewards critical questioning of past results and of colleagues.
Most technologies benefit from a solid research foundation in academia
and ample opportunity for peer-to-peer stimulation and critical assessment,
review and critique through conferences, seminars, publishing, and more.
These elements provide a rich set of paths through which new ideas and
skepticism can travel and opportunities for scientists to step away from
their day-to-day work and take a longer-term view. The scientific culture
encourages cautious, precise statements and discourages statements that go
beyond established facts; it is acceptable for colleagues to challenge one an-
other, even if the challenger is more junior. The forensic science disciplines
will profit enormously by full adoption of this scientific culture.
CONCLUSION
The way in which science is conducted is distinct from, and comple-
mentary to, other modes by which humans investigate and create. The
methods of science have a long history of successfully building useful and
trustworthy knowledge and filling gaps while also correcting past errors.
The premium that science places on precision, objectivity, critical thinking,
careful observation and practice, repeatability, uncertainty management,
and peer review enables the reliable collection, measurement, and interpre-
tation of clues in order to produce knowledge.
The effects of contextual top-down processing on matching fingerprints. Journal of Applied
Cognitive Psychology 19:799-809; and B. Schiffer and C. Champod. 2007. The potential
(negative) influence of observational biases at the analysis stage of fingerprint individualiza-
tion. Forensic Science International 167:116-120.
5
Descriptions of Some Forensic
Science Disciplines
This chapter describes the methods of some of the major forensic
science disciplines. It focuses on those that are used most commonly for
investigations and trials as well as on those that have been cause for con-
cern in court or elsewhere because their reliability has not been sufficiently
established in a systematic (scientific) manner in accordance with the prin-
ciples discussed in Chapter 4. The chapter focuses primarily on the forensic
science disciplines’ capability for providing evidence that can be presented
in court. As such, there is considerable discussion about the reliability and
precision of results—attributes that factor into probative value and admis-
sibility decisions. It should be recalled, however, that forensic science also
provides great value to law enforcement investigations, and even those
forensic science disciplines whose scientific foundation is currently limited
might have the capacity (or the potential) to provide probative informa-
tion to advance a criminal investigation.1 This chapter also provides the
committee’s summary assessment of each of these disciplines.2
1 For example, forensic odontology might not be sufficiently grounded in science to be ad-
missible under Daubert, but this discipline might be able to reliably exclude a suspect, thereby
enabling law enforcement to focus its efforts on other suspects. And forensic science methods
that do not meet the standards of admissible evidence might still offer leads to advance an
investigation.
2 The chapter does not discuss eyewitness identification or line-ups, because these techniques
do not normally rely on forensic scientists for analysis or implementation. They clearly are of
major importance for investigations and trials, and their effective use and interpretation relies
on scientific knowledge and continuing research. For similar reasons, this chapter does not
delve into the polygraph. The validity of polygraph testing for security screening was addressed
127
128
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
Because forensic science aims to glean information from a wide variety
of clues and evidence associated with a crime, it deals with a broad range
of tools and with evidence of highly variable quality. In general, the forensic
science disciplines are pragmatic, with practitioners adopting, adapting, or
developing whatever tools and technological aids they can to distill useful
information from crime scene evidence. Many forensic science methods
have been developed in response to such evidence—combining experience-
based knowledge with whatever relevant science base exists in order to
create a procedure that returns useful information. Although some of the
techniques used by the forensic science disciplines—such as DNA analysis,
serology, forensic pathology, toxicology, chemical analysis, and digital and
multimedia forensics—are built on solid bases of theory and research, many
other techniques have been developed heuristically. That is, they are based
on observation, experience, and reasoning without an underlying scientific
theory, experiments designed to test the uncertainties and reliability of the
method, or sufficient data that are collected and analyzed scientifically.
In the course of its deliberations, the committee received testimony
from experts in many forensic science disciplines concerning current prac-
tices, validity, reliability and errors, standards, and research.3 From this
testimony and from many written submissions, as well as from the personal
experiences of the committee members, the committee developed the con-
sensus views presented in this chapter.
BIOLOGICAL EVIDENCE
Biological evidence is provided by specimens of a biological origin that
are available in a forensic investigation. Such specimens may be found at the
scene of a crime or on a person, clothing, or weapon. Some—for example,
pet hairs, insects, seeds, or other botanical remnants—come from the crime
scene or from an environment through which a victim or suspect has re-
cently traversed. Other biological evidence comes from specimens obtained
directly from the victim or suspect, such as blood, semen, saliva, vaginal
secretions, sweat, epithelial cells, vomitus, feces, urine, hair, tissue, bones,
and microbiological and viral agents. The most common types of biological
evidence collected for examination are blood, semen, and saliva. Human
biological evidence that contains nuclear DNA can be particularly valuable
because the possibility exists to associate that evidence with one individual
with a degree of reliability that is acceptable for criminal justice.
in National Research Council, Committee to Review the Scientific Evidence on the Polygraph.
2003. The Polygraph and Lie Detection. Washington, DC: The National Academies Press. It
does not cover forensic pathology, because that field is addressed in Chapter 9.
A complete list of those who provided testimony to the committee is included in Appendix B.
FORENSIC SCIENCE DISCIPLINES
129
Sample Data and Collection
At the crime scene, biological evidence is located, documented, col-
lected, and preserved for subsequent analysis in the crime laboratory. Lo-
cating and recognizing biological evidence can be more difficult than a
layperson would presume. For example, blood is not always red, some red
substances are not blood, and most biological evidence, such as saliva or
semen, is not readily visible. Crime scene investigators locate biological
evidence through tests that screen for the presence of a particular bio-
logical fluid (e.g., blood, semen, saliva), and investigators have a choice of
techniques.4 For blood they might use an alternate light source (ALS) at
415nm, the wavelength under which bloodstains absorb light and are thus
more visible to the naked eye. Most commonly, though, the screening test
for blood is a catalytic chemical test that turns color or luminesces in the
presence of blood. Scene investigators may also use Luminol, fluorescein,
or crystal violet to identify areas at the scene where attempts were made to
clean a bloody crime scene.
These tests for blood may also locate other evidence that should be
collected and taken to the laboratory for analysis. Recently, immunological
tests that can identify human hemoglobin or glycophorin A have become
available. These are blood-specific proteins that can be demonstrated to be
of human origin. At some point in the future, these immunological tests
may replace standard chemical tests, and, although more expensive, they
are more specific because they identify blood conclusively instead of just
presumptively. Investigators also have several techniques for locating semen
at the crime scene. Commonly they rely on an ALS, under which semen,
other biological fluids, and some other evidence will luminesce. More re-
cently, immunological tests can be used to identify seminal plasma proteins,
for example, prostate specific antigen (p30 or PSA) or semenogelin.5
Finding saliva at the scene is mostly happenstance. Although it lumi-
nesces with the ALS at specific wavelengths, the glow is not as strong, and
a weaker ALS light source may not highlight it well and possibly not at
all. Thus, it can be easily missed. Screening tests for saliva are chemical
tests that identify amylase, an enzyme occurring in high concentrations in
saliva. But the screening is not definitive, because other types of tissue also
4 Interpreting the results of any screening test requires expertise and experience. Many crime
scene investigators have the requisite experience, but they may lack a scientific background,
and it is not always straightforward to correctly interpret the results of screening tests. Crime
scene investigations that require science-based screening tools are most reliable if someone is
involved who understands the physics and chemistry of those tools.
5 I. Sato, M. Sagi, A. Ishiwari, H. Nishijima, E. Ito, and T. Mukai. 2002. Use of the
“SMITEST” PSA card to identify the presence of prostate-specific antigen in semen and male
urine. Forensic Science International 127(1-2):71-74.
130
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
contain amylase, including the particular type (AMY 1) that is associated
with saliva.
Analyses
Although the forensic use of nuclear DNA is barely 20 years old, DNA
typing is now universally recognized as the standard against which many
other forensic individualization techniques are judged. DNA enjoys this
preeminent position because of its reliability and the fact that, absent fraud
or an error in labeling or handling, the probabilities of a false positive are
quantifiable and often miniscule. However, even a very small (but nonzero)
probability of false positive can affect the odds that a suspect is the source
of a sample with a matching DNA profile.6 The scientific bases and reli-
ability of other types of biological analysis are also well established, but
absent nuclear DNA, they can only narrow the field of suspects, not suggest
any particular individual.
Testing biological evidence in the laboratory involves the use of a logi-
cal sequence of analyses designed to identify what a substance is and then
from whom it came. The sequence begins with a forensic biologist locat-
ing the substance on the evidence. This is followed by a presumptive test
that would give more information about the substance, typically using the
same tests employed by scene investigators: the ALS, enzymatic, chemical,
or immunological tests. Once the material (e.g., blood, semen, or saliva) is
known, an immunological test or a human DNA test is run to determine
whether the sample comes from a human or an animal.
The final step in the analytical sequence procedure is to identify the
source of the biological material. If a sufficient sample is present and is
probative, the forensic biologist prepares the material for DNA testing. The
analyst who conducts the DNA test may or may not be the same person
who examines the original physical evidence, depending on laboratory
policies.
A decision might be required regarding the type of DNA testing to
employ. Two primary types of DNA tests are conducted in U.S. forensic
laboratories: nuclear testing and mitochondrial DNA (mtDNA) testing,
with several variations of the former. For most biological evidence having
evidentiary significance, forensic DNA laboratories employ nuclear test-
ing routinely,7 and testing for the 13 core Short Tandem Repeat (STR)
6 W.C. Thompson, F. Taroni, and C.G.G. Aitken. 2003. How the probability of a false posi-
tive affects the value of DNA evidence. Journal of Forensic Sciences 48(1):47-54.
7 T.R. Moretti, A.L. Baumstark, D.A. Defenbaugh, K.M. Keys, J.B. Smerick, and B. Budowle
B. 2001. Validation of short tandem repeats (STRs) for forensic usage: Performance testing of
fluorescent multiplex STR systems and analysis of authentic and simulated forensic samples.
Journal of Forensic Sciences 46(3):647-660.
FORENSIC SCIENCE DISCIPLINES
131
polymorphisms is the first line of attack.8 The results are entered into the
Federal Bureau of Investigation’s (FBI’s) Combined DNA Indexing System
(CODIS) and are searched against DNA profiles already in one of three
databases: a convicted felon database, a forensic database containing
DNA profiles from crime scenes, and a database of DNA from unidenti-
fied persons.
Sometimes the evidence dictates testing just for Y STRs, which assesses
only the Y (male) chromosome. In sexual assaults for which only small
amounts of male nuclear DNA are available (e.g., a large excess of vaginal
DNA), it is possible to obtain a Y STR profile of the male who left the se-
men. Unlike the 13 core loci used in CODIS searches, where a match of all
13 is a strong indicator that both samples come from the same individual, Y
STR testing is not as definitive with respect to identifying a single person. A
third nuclear test involves the analysis of single nucleotide polymorphisms
(SNPs). Although no public forensic DNA laboratory in the United States
is routinely analyzing forensic evidence for SNPs, the utility of this genomic
information for cases in which the DNA is too damaged to allow standard
testing has garnered attention since its use in the World Trade Center iden-
tification effort.9
If insufficient nuclear DNA is present for STR testing, or if the exist-
ing nuclear DNA is degraded, two options potentially are available. One
technique amplifies the amount of DNA available, although this technique
is not widely available in U.S. forensic laboratories. A second alternative is
to sequence mitochondrial DNA (mtDNA). Since 1996, it has been possible
to compare single-source crime scene samples and samples from the victim
or defendant on the basis of mtDNA. Four FBI-supported mtDNA labo-
ratories and a few private mtDNA laboratories conduct DNA casework.
This technique has been particularly helpful with regard to hairs—which do
not contain enough nuclear DNA to enable analysis with current methods
unless the root is present—and bones and teeth. Because it measures only
a single locus of the genome, mtDNA analysis is much less discriminating
than nuclear DNA analysis; all people with a common female ancestor
(within the past few generations) share a common profile. But mtDNA
testing has forensic value in its ability to include or exclude an individual
as its source.
Laboratories entering the results of forensic DNA testing into CODIS
must meet specific quality guidelines, which include the requirement that
8 Some laboratories are now using 16 loci, 13 of which are the original core loci.
9 B. Leclair, R. Shaler, G.R. Carmody, K. Eliason, B.C.Hendrickson, T. Judkins, M.J. Norton,
C. Sears, and T. Scholl. 2007. Bioinformatics and human identification in mass fatality in-
cidents: The World Trade Center disaster. Journal of Forensic Sciences 52(4):806-819. Epub
May 25, 2007.
132
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
the laboratory be accredited and that specific procedures be in place and
followed. In accredited laboratories, forensic DNA personnel must take
proficiency tests and must meet specific educational and training require-
ments. (See Chapter 8 for further discussion.) Laboratory analyses are
conducted by scientists with degrees ranging from a bachelor’s degree in
science to a doctoral degree. Each forensic DNA laboratory has a techni-
cal leader, who normally must meet additional experience and educational
requirements.
Although DNA laboratories are expected to conduct their examina-
tions under stringent quality controlled environments, errors do occa-
sionally occur. They usually involve situations in which interpretational
ambiguities occur or in which samples were inappropriately processed
and/or contaminated in the laboratory. Errors also can occur when there
are limited amounts of DNA, which limits the amount of test information
and increases the chance of misinterpretation. Casework reviews of mtDNA
analysis suggest a wide range in the quality of testing results that include
contamination, inexperience in interpreting mixtures, and differences in
how a test is conducted.10
Reporting of Results
FBI quality guidelines require that reports from forensic DNA analysis
must contain, at a minimum, a description of the evidence examined, a list-
ing of the loci analyzed, a description of the methodology, results and/or
conclusions, and an interpretative statement (either quantitative or qualita-
tive) concerning the inference to be drawn from the analysis.11
10 Personal communication, Terry Melton, Mitotyping Laboratory. December 2007. See
also L. Prieto; A. Alonso; C. Alves; M. Crespillo; M. Montesino; A. Picornell; A. Brehm; J.L.
Ramirez; M.R. Whittle; M.J. Anjos; I. Boschi; J. Buj; M. Cerezo; S. Cardoso; R. Cicarelli; D.
Comas; D. Corach; C. Doutremepuich; R.M. Espinheira; I. Fernandez-Fernandez; S. Filippini;
Julia Garcia-Hirschfeld; A. Gonzalez; B. Heinrichs; A. Hernandez; F.P.N. Leite; R.P. Lizarazo;
A.M. Lopez-Parra; M. Lopez-Soto; J.A. Lorente; B. Mechoso; I. Navarro; S. Pagano; J.J.
Pestano; J. Puente; E. Raimondi; A. Rodriguez-Quesada; M.F. Terra-Pinheiro; L. Vidal-Rioja;
C. Vullo; A. Salas. 2008. GEP-ISFG collaborative exercise on mtDNA: Reflections about
interpretation, artefacts and DNA mixtures. Forensic Science International: Genetics 2(2):126-
133; and A. Salas, L. Prieto, M. Montesino, C. Albarrán, E. Arroyo, M. Paredes-Herrera, A.
Di Lonardo, C. Doutremepuich, I. Fernández-Fernández, A. de la Vega. 2005. Mitochondrial
DNA error prophylaxis: Assessing the causes of errors in the GEP’02-03 proficiency testing
trial. Forensic Science International 148(2-3):191-198.
11 DNA Advisory Board. 2000. Quality assurance standards for forensic DNA test-
com/conference04/TWGDAM/Quality_Assurance_Standards_2.pdf.
FORENSIC SCIENCE DISCIPLINES
133
Summary Assessment
Unlike many forensic techniques that were developed empirically within
the forensic science community, with limited foundation in scientific theory
or analysis, DNA analysis is a fortuitous by-product of cutting-edge sci-
ence. Eminent scientists contributed their expertise to ensuring that DNA
evidence offered in a courtroom would be valid and reliable (e.g., in the
1989 New York case, People v. Castro), and by 1996 the National Academy
of Sciences had convened two committees that issued influential recom-
mendations on handling DNA forensic science.12 As a result, principles
of statistics and population genetics that pertain to DNA evidence were
clarified, the methods for conducting DNA analyses and declaring a match
became less subjective, and quality assurance and quality control protocols
were designed to improve laboratory performance.
DNA analysis is scientifically sound for several reasons: (1) there are
biological explanations for individual-specific findings; (2) the 13 STR loci
used to compare DNA samples were selected so that the chance of two
different people matching on all of them would be extremely small; (3)
the probabilities of false positives have been explored and quantified in
some settings (even if only approximately); (4) the laboratory procedures
are well specified and subject to validation and proficiency testing; and (5)
there are clear and repeatable standards for analysis, interpretation, and
reporting. DNA analysis also has been subjected to more scrutiny than
any other forensic science discipline, with rigorous experimentation and
validation performed prior to its use in forensic investigations. As a result
of these characteristics, the probative power of DNA is high. Of course,
DNA evidence is not available in every criminal investigation, and it is still
subject to errors in handling that can invalidate the analysis. In such cases,
other forensic techniques must be applied. The probative power of these
other methods can be high, alone or in combination with other evidence.
This power likely can be improved by strengthening the methods’ scientific
foundations and practice, as has occurred with forensic DNA analysis.
ANALYSIS OF CONTROLLED SUBSTANCES
The term “illicit drugs” is widely used to describe abused substances.
Other terms that are used include “abused drugs,” “illegal drugs,” “street
drugs,” and, in the United States, “controlled substances.” The latter term
refers specifically to drugs that are controlled by federal and state laws.13
12 National Research Council. 1992. DNA Technology in Forensic Science. Washington,
DC: National Academy Press; National Research Council. 1996. The Evaluation of Forensic
DNA Evidence: An Update. Washington, DC: National Academy Press.
13 See, e.g., 21 U.S.C.A. § 802(6).
134
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
The analysis of controlled substances is a mature forensic science dis-
cipline and one of the areas with a strong scientific underpinning. The ana-
lytical methods used have been adopted from classical analytical chemistry,
and there is broad agreement nationwide about best practices.14 In 1997,
the U.S. Drug Enforcement Administration and the Office of National Drug
Control Policy co-sponsored the formation of the Technical Working Group
for the Analysis of Seized Drugs, now known as the Scientific Working
Group for the Analysis of Seized Drugs (SWGDRUG). This organization
brings together more than 20 forensic practitioners from all over the world
to develop standards for the analysis and reporting of illicit drug cases.
Their standards are being widely adopted by drug analysis laboratories in
the United States and worldwide.
Sample Data and Collection
Controlled substances typically are seized by police officers, narcotics
agents, and detectives through undercover buys, raids on drug houses and
clandestine drug laboratories, and seizures on the streets. In some cases, fo-
rensic chemists are sent to clandestine laboratory operations to help render
the laboratory safe and help with evidence collection. The seized drugs may
be in the form of powders or adulterated powders, chunks of smokeable or
injectable material, legitimate and clandestine tablets and capsules, or plant
materials or plant extracts.
Analyses
Controlled substances are analyzed by well-accepted standard schemes
or protocols. Few drug chemists have the requisite botanical background
to identify any common illicit plants other than marijuana; thus, in cases
that require botanical identification, the assistance of outside experts is
enlisted.
Sampling can be a major issue in the analysis of controlled substances.
Although sometimes only trace amounts of a drug are present (e.g., in a sy-
ringe used to inject heroin), at other times there are hundreds or thousands
of packages of drugs or very large bags or bales. SWGDRUG and others
have proposed statistical and nonstatistical methods for sampling,15 and a
wide variety of methods are used.
Most controlled substances are subjected first to a field test for pre-
14 See F. Smith and J.A. Siegel (eds.). 2004. Handbook of Forensic Drug Analysis. Burling-
ton, MA: Academic Press.
15 Scientific Working Group for the Analysis of Seized Drugs (SWGDRUG) Recommenda-
tions. Available at www.swgdrug.org/approved.htm.
FORENSIC SCIENCE DISCIPLINES
135
sumptive identification. This is followed by gas chromatography-mass spec-
trometry (GC-MS), in which chromatography separates the drug from any
diluents or excipients, and then mass spectrometry is used to identify the
drug. This is the near universal test for identifying unknown substances.
Marijuana is an exception, because it is identified normally through a se-
quence of tests—a presumptive color test, followed by low-powered micro-
scopic identification, and finally by thin-layer chromatography.
Reporting of Results
Most drug chemists produce terse reports for attorneys and courts. The
reports contain administrative data and a short description of the evidence.
The weight or number of exhibits is stated and then the results of the analy-
sis. A typical report for a marijuana case might read as follows:
Received: Item 1—a sealed plastic bag containing 25.6 g of green-
brown plant material.
Results:
The green-brown plant material in item 1 was identified as
marijuana.
Some laboratories might mention the tests that were conducted, but in
most cases the spectra, chromatograms, and other evidence of the analysis
and the chemist’s notes are not submitted. Likewise, possible sources of
error and statistical data are not commonly included. From a scientific
perspective, this style of reporting is often inadequate, because it may not
provide enough detail to enable a peer or other courtroom participant to
understand and, if needed, question the sampling scheme, process(es) of
analysis, or interpretation.
Summary Assessment
The chemical foundations for the analysis of controlled substances are
sound, and there exists an adequate understanding of the uncertainties and
potential errors. SWGDRUG has established a fairly complete set of recom-
mended practices.16 It also provides pointers to a number of guidelines for
statistical sampling, both for illegal drugs per se (created by the European
Network of Forensic Science Institutes) and for materials more generally
(created by the American Society for Testing and Materials).
The SWGDRUG recommendations include a menu of analytical chem-
istry techniques that are considered acceptable in certain circumstances.
Because this menu was constructed to be applicable worldwide, it includes
136
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
options that allow laboratories to substitute a concatenation of simple
methods if they do not have access to the preferred analytical equipment
(e.g., GC-MS). It is questionable, however, whether all of the possible com-
binations recommended by SWGDRUG would be acceptable in a scientific
sense, if one’s goal were to identify and classify a completely unknown
substance. The committee has been told that experienced forensic chemists
and good forensic laboratories understand which tests (or combinations of
tests) provide adequate reliability, but the SWGDRUG recommendations
do not ensure that these tests will be used. This ambiguity would be a less
significant issue if the reports presented in court contained sufficient detail
about the methods of analysis.
FRICTION RIDGE ANALYSIS
Fingerprints, palm prints, and sole prints have been used to identify
people for more than a century in the United States. Collectively, the analy-
sis of these prints is known as “friction ridge analysis,” which consists of
experience-based comparisons of the impressions left by the ridge structures
of volar (hands and feet) surfaces. Friction ridge analysis is an example of
what the forensic science community uses as a method for assessing “indi-
vidualization”—the conclusion that a piece of evidence (here, a pattern left
by friction ridges) comes from a single unambiguous source. Friction ridge
analysis shares similarities with other experience-based methods of pattern
recognition, such as those for footwear and tire impressions, toolmarks,
and handwriting analysis, all of which are discussed separately below.
Friction ridge analysis is performed in various settings, including ac-
credited crime laboratories and nonaccredited facilities. Nonaccredited
facilities may be crime laboratories, police “identification units,” or private
practice (consultants). In some instances, the latent print examiner is em-
ployed solely to perform latent print casework. Some examiners may also
perform other types of forensic casework (e.g., footwear and tire impres-
sions, firearms analysis). In some agencies, fingerprint examiners also are
required to respond to crime scenes and can be sworn officers who also
perform police officer/detective duties.
The training of personnel to perform latent print identifications varies
from agency to agency. Agencies may have a formalized training program,
may use an informal mentoring process, or may send new examiners to
a one- to two-week course. The International Association for Identifica-
tion (IAI) offers a training publication, “Friction Ridge Skin Identification
Training Manual,”17 and the Scientific Working Group on Friction Ridge
17 International Association for Identification. Friction Ridge Skin Identification Training
Manual. Available at www.theiai.org.
FORENSIC SCIENCE DISCIPLINES
137
Analysis, Study and Technology (SWGFAST) offers a guideline, “Training
to Competency for Latent Print Examiners.”18 Although these are excellent
resources, they are not required, and there is no auditing of the content of
training programs developed by nonaccredited agencies. The IAI also of-
fers a certification test that measures both the knowledge and skill of latent
print examiners; however, not all agencies require latent print examiners to
achieve and maintain certification.
Method of Data Collection and Analysis
The technique used to examine prints made by friction ridge skin is
described by the acronym ACE-V: “Analysis, Comparison, Evaluation, and
Verification.”19 It has been described in forensic literature as a means of
comparative analysis of evidence since 1959.20 The process begins with the
analysis of the unknown friction ridge print (now often a digital image of
a latent print). Many factors affect the quality and quantity of detail in the
latent print and also introduce variability in the resulting impression. The
examiner must consider the following:
(1) Condition of the skin—natural ridge structure (robustness of the
ridge structure), consequences of aging, superficial damage to the
skin, permanent scars, skin diseases, and masking attempts.
(2) Type of residue—natural residue (sweat residue, oily residue, com-
binations of sweat and oil); other types of residue (blood, paint,
etc.); amount of residue (heavy, medium, or light); and where the
residue accumulates (top of the ridge, both edges of the ridge, one
edge of the ridge, or in the furrows).
(3) Mechanics of touch—underlying structures of the hands and feet
(bone creates areas of high pressure on the surface of the skin);
flexibility of the ridges, furrows, and creases; the distance adja-
cent ridges can be pushed together or pulled apart during lateral
movement; the distance the length of a ridge might be compressed
or stretched; the rotation of ridge systems during torsion; and the
effect of ridge flow on these factors.
(4) Nature of the surface touched—texture (rough or smooth), flex-
ibility (rigid or pliable), shape (flat or curved), condition (clean or
dirty), and background colors and patterns.
SWGFAST.org.
19 Ashbaugh, op. cit.; Triplette and Cooney, op. cit.; J. Vanderkolk. 2004. ACE-V: A model.
Journal of Forensic Identification 54(1):45-52; SWGFAST. 2002. Friction Ridge Examination
Methodology for Latent Print Examiners. Available at www.SWGFAST.org.
20 R.A. Huber. 1959-1960. Expert witness. Criminal Law Quarterly 2:276-296.
138
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
(5) Development technique—chemical signature of the technique and
consistency of the chemical signature across the impression.
(6) Capture technique—photograph (digital or film) or lifting material
(e.g., tape or gelatin lifter).
(7) Size of the latent print or the percentage of the surface that is avail-
able for comparison.
The examiner also must perform an analysis of the known prints (taken
from a suspect or retrieved from a database of fingerprints), because many
of the same factors that affect the quality of the latent print can also affect
the known prints.
If the latent print does not have sufficient detail for either identification
or exclusion, it does not undergo the remainder of the process (comparison
and evaluation). These insufficient prints are often called “of no value” or
“not suitable” for comparison. Poor-quality known prints also will end
the examination. If the examiner deems that there is sufficient detail in the
latent print (and the known prints), the comparison of the latent print to
the known prints begins.
Visual comparison consists of discerning, visually “measuring,” and
comparing—within the comparable areas of the latent print and the known
prints—the details that correspond. The amount of friction ridge detail
available for this step depends on the clarity of the two impressions. The
details observed might include the overall shape of the latent print, ana-
tomical aspects, ridge flows, ridge counts, shape of the core, delta location
and shape, lengths of the ridges, minutia location and type, thickness of the
ridges and furrows, shapes of the ridges, pore position, crease patterns and
shapes, scar shapes, and temporary feature shapes (e.g., a wart).
At the completion of the comparison, the examiner performs an evalua-
tion of the agreement of the friction ridge formations in the two prints and
evaluates the sufficiency of the detail present to establish an identification
(source determination).21 Source determination is made when the examiner
concludes, based on his or her experience, that sufficient quantity and qual-
ity of friction ridge detail is in agreement between the latent print and the
known print. Source exclusion is made when the process indicates sufficient
disagreement between the latent print and known print. If neither an iden-
tification nor an exclusion can be reached, the result of the comparison is
inconclusive. Verification occurs when another qualified examiner repeats
the observations and comes to the same conclusion, although the second
examiner may be aware of the conclusion of the first. A more complete de-
21 Ashbaugh, op. cit.; SWGFAST. 2002. Friction Ridge Examination Methodology for
Latent Print Examiners.
FORENSIC SCIENCE DISCIPLINES
139
scription of the steps of ACE-V and an analysis of its limitations is provided
in a paper by Haber and Haber.22
Although some Automated Fingerprint Identification Systems (AFIS)
permit fully automated identification of fingerprint records related to crimi-
nal history (e.g., for screening job applicants), the assessment of latent
prints from crime scenes is based largely on human interpretation. Note
that the ACE-V method does not specify particular measurements or a
standard test protocol, and examiners must make subjective assessments
throughout. In the United States, the threshold for making a source iden-
tification is deliberately kept subjective, so that the examiner can take into
account both the quantity and quality of comparable details. As a result,
the outcome of a friction ridge analysis is not necessarily repeatable from
examiner to examiner. In fact, recent research by Dror23 has shown that
experienced examiners do not necessarily agree with even their own past
conclusions when the examination is presented in a different context some
time later.
This subjectivity is intrinsic to friction ridge analysis, as can be seen
when comparing it with DNA analysis. For the latter, 13 specific segments
of DNA (generally) are compared for each of two DNA samples. Each of
these segments consists of ordered sequences of the base pairs, called A, G,
C, and T. Studies have been conducted to determine the range of variation
in the sequence of base pairs at each of the 13 loci and also to determine
how much variation exists in different populations. From these data, sci-
entists can calculate the probability that two DNA samples from different
people will have the same permutations at each of the 13 loci.
By contrast, before examining two fingerprints, one cannot say a priori
which features should be compared. Features are selected during the com-
parison phase of ACE-V, when a fingerprint examiner identifies which
features are common to the two impressions and are clear enough to be
evaluated. Because a feature that was helpful during a previous compari-
son might not exist on these prints or might not have been captured in the
latent impression, the process does not allow one to stipulate specific mea-
surements in advance, as is done for a DNA analysis. Moreover, a small
stretching of distance between two fingerprint features, or a twisting of
angles, can result from either a difference between the fingers that left the
prints or from distortions from the impression process. For these reasons,
population statistics for fingerprints have not been developed, and friction
ridge analysis relies on subjective judgments by the examiner. Little research
22 L. Haber and R.N. Haber. 2008. Scientific validation of fingerprint evidence under
Daubert. Law, Probability, and Risk 7(2):87-109.
23 I.E. Dror and D. Charlton. 2006. Why experts make errors. Journal of Forensic Identi-
fication 56(4):600-616.
140
STRENGTHENING FORENSIC SCIENCE IN THE UNITED STATES
has been directed toward developing population statistics, although more
would be feasible.24
Methods of Interpretation
The determination of an exclusion can be straightforward if the exam-
iner finds detail in the latent print that does not match the corresponding
part of the known print, although distortions or poor image quality can
complicate this determination. But the criteria for identification are much
harder to define, because they depend on an examiner’s ability to discern
patterns (possibly complex) among myriad features and on the examiner’s
experience judging the discriminatory value in those patterns. The clarity of
the prints being compared is a major underlying factor. For 10-print finger-
print cards, which tend to have good clarity, even automated pattern-recog-
nition software (which is not as capable as human examiners) is successful
enough in retrieving matching sets from databases to enjoy widespread use.
When dealing with a single latent print, however, the interpretation task
becomes more challenging and relies more on the judgment of the examiner.
The committee heard presentations from friction ridge experts who assured
it that friction ridge identification works well when a careful examiner
works with good-quality latent prints. Clearly, the reliability of the ACE-V
process could be improved if specific measurement criteria were defined.
Those criteria become increasingly important when working with latent
prints that are smudged and incomplete, or when comparing impressions
from two individuals whose prints are unusually similar.
The fingerprint community continues to assert that the ability to see
latent print detail is an acquired skill attained only through repeated expo-
sure to friction ridge impressions. In their view, a lengthy apprenticeship
(typically two years, at the FBI Laboratory) with an experienced latent print
examiner enables a new examiner to develop a sense of the rarity of features
and groups of features; the rarity of particular kinds of ridge flows; the
frequency of features in different areas of the hands and feet; the degree to
which differences can be accounted for by mechanical distortion of the skin;
a sense of how to extract detail from background noise; and a sense of how
much friction ridge detail could be common to two prints from different
24 See, e.g., E. Gutiérrez, V. Galera, J.M. Martínez, and C. Alonso. 2007. Biological vari-
ability of the minutiae in the fingerprints of a sample of the Spanish population. Forensic
Science International 172(2-3):98-105. For information about the basic availability of data,
see C. Champod, C.J. Lennard, P.A. Margot, and M. Stoilovic. 2004. Fingerprints and other
ridge skin impressions. Boca Raton, FL: CRC Press; D.A. Stoney. 2001. “Measurement of
Fingerprint Individuality.” In: H.C. Lee and R.E. Gaensslen (eds.). Advances in Fingerprint
Technology. 2nd ed. Boca Raton, FL: CRC Press; pp. 327-387.
|
||
|
|
|