Software Engineering Empirical Research Radar

Benchmark Map

Common software-engineering datasets and whether SEER Radar has local evidence or a local copy.

Status meanings: dataset-copy-found means a local directory appears to contain the benchmark; papers-or-notes-found means local papers, abstracts, fulltexts, or notes mention it; not-found-locally means no local evidence was found yet.

BenchmarkCategoryLocal StatusLocal Evidence
TravisTorrent
A Travis CI and GitHub dataset frequently used for CI build, test, and failure prediction studies.
Continuous integration build historypapers-or-notes-foundPDFs: 4 · abstracts: 5 · fulltexts: 4 · notes: 18
Defects4J
A widely used benchmark of reproducible Java bugs with buggy/fixed versions and developer tests.
Real Java bugs / testing / fault localizationdataset-copy-foundPDFs: 2 · abstracts: 2 · fulltexts: 2 · notes: 28
SWE-bench
A benchmark of real GitHub issues where systems must modify repositories to resolve software problems.
LLM agents / real GitHub issue resolutionpapers-or-notes-foundPDFs: 2 · abstracts: 0 · fulltexts: 2 · notes: 2
BuildSheriff Dataset
CI test-failure triage data associated with change-aware clustering of broken builds.
CI failure triagepapers-or-notes-foundPDFs: 1 · abstracts: 1 · fulltexts: 1 · notes: 0
iFixFlakies
A framework and dataset direction around detecting and fixing order-dependent flaky tests.
Order-dependent flaky testspapers-or-notes-foundPDFs: 1 · abstracts: 1 · fulltexts: 1 · notes: 1
IDoFT
A dataset of flaky tests collected from open-source projects.
Flaky testspapers-or-notes-foundPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 2
SIR / Siemens
Classic benchmark suites used heavily in regression testing, test prioritization, and fault localization.
Classic regression testing and fault localizationpapers-or-notes-foundPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 12
Bears
A benchmark of bugs collected from continuous integration failures in Java projects.
Java bugs from CInot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0
BugSwarm
A dataset of reproducible CI failures and fixes packaged for empirical software engineering research.
CI failures / reproducible build-test failuresnot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0
Bugs.jar
A benchmark of real Java bugs from large open-source projects.
Java bugs / regression testingnot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0
Codeflaws
A benchmark of C programming faults derived from programming contest submissions.
Program repair / C programming contest bugsnot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0
ManyBugs / IntroClass
Classic C benchmark suites for automated program repair and fault localization.
Program repairnot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0
QuixBugs
A small benchmark of buggy programs used in program repair and debugging studies.
Program repair / small algorithmic bugsnot-found-locallyPDFs: 0 · abstracts: 0 · fulltexts: 0 · notes: 0