Independent testing for research teams

Choose the research question you need answered. We scope the evaluation around the claim, available evidence, and intended use.

For research

Start with the uncertainty, not a service label.

Open the question closest to your situation. A fit check determines the final scope.

Three computational researchers compare a manuscript, repository output, and a computing test setup.
Reviewing the paper, artifact, execution record, and measurement together.
Research reproduction and replicationRerun a reported result from the supplied materials, or test the same claim with an independent implementation, dataset, or experimental design.

Useful when

  • A reported result needs independent confirmation.
  • A journal, conference, or funder needs more than an artifact check.
  • You need to distinguish rerunning the original work from testing the claim independently.

The question

Can we reproduce the reported result, and does the claim hold in an independent replication?

What we test

  • Reported methods and artifacts
  • Computational environment and execution
  • Results and declared tolerances
  • Independent replication design

What you receive

  • Reproduction record
  • Replication plan or results
  • Deviations and failure boundaries
  • Conclusion and open questions

What to send

  • Paper, preprint, or claim specification
  • Available code, data, environment, and execution instructions
  • Expected outputs and acceptable tolerances

A fit check determines whether the work should cover reproduction, replication, or both. Scope depends on the claim, available materials, access, and the amount of independent reconstruction required.

For: Researchers, journals, conferences, funders, and public-interest organizations

Discuss a reproduction or replication
Claim and method reviewAssess whether the method, analysis, and evidence support the stated claim.

Useful when

  • A team wants an external review before submission or release.
  • A central claim depends on assumptions, analysis choices, or indirect evidence.
  • Reviewers need a traceable account of how the evidence supports the claim.

The question

Do the method, analysis, and evidence support the stated claim?

What we test

  • Claim specification and scope
  • Research design and assumptions
  • Measurement and statistical analysis
  • Limitations and validity threats

What you receive

  • Claim-to-evidence map
  • Method review
  • Material evidence gaps
  • Recommended revisions

What to send

  • Draft or published paper
  • Analysis code, data, and supporting documentation
  • The claims or reviewer concerns that need attention

This is an advisory review. If VRL later evaluates the same work independently, we disclose the earlier role and assign different reviewers.

For: Researchers, research teams, laboratories, journals, and conferences

Request a claim review
Artifact reviewCheck whether code, data, environments, and instructions can be retrieved, understood, and executed independently.

Useful when

  • Code and data must be usable by reviewers or future researchers.
  • A team is preparing an artifact for submission, release, or archival.
  • Execution currently depends on undocumented knowledge from the original team.

The question

Can an external evaluator inspect and execute the work from the supplied materials?

What we test

  • Documentation and expected outputs
  • Dependencies and automation
  • Data access and licensing
  • Versioning, archival, and hosting readiness

What you receive

  • Blind-first reconstruction log
  • Assistance record
  • Prioritized fixes
  • Preservation and hosting plan

What to send

  • Repository or archive
  • Data, environment files, and access instructions
  • Expected outputs and any known limitations

The review begins as an external evaluator would: with the supplied materials. Assistance from the team is recorded so the final report distinguishes documented steps from guided reconstruction.

For: Researchers, laboratories, journals, conferences, and artifact evaluators

Request an artifact review
Benchmark auditTest baselines, workloads, measurements, statistics, and claimed advantages.

Useful when

  • A performance claim depends on baseline choice or tuning.
  • Workloads, resource limits, or cost assumptions may change the comparison.
  • A benchmark result will influence publication, procurement, or adoption.

The question

Does the claimed advantage survive fair baselines, representative workloads, and resource parity?

What we test

  • Baseline choice and tuning
  • Workload and resource parity
  • Measurement and statistics
  • Cost and operating conditions

What you receive

  • Fair-baseline reconstruction
  • Raw measurements and analysis
  • Sensitivity results
  • Failure boundaries

What to send

  • Benchmark code, workloads, and reported results
  • Baseline implementations and tuning instructions
  • Hardware, cost, and operating assumptions

The audit focuses on the comparison behind the claim. We agree on baselines, workloads, parity conditions, and material sensitivity tests before execution.

For: Researchers, laboratories, standards groups, technical buyers, and maintainers

Request a benchmark audit
Independent verification reportPublish a versioned record of what reproduced, what replicated, what failed, and what remains uncertain.

Useful when

  • A claim needs a public, citable evaluation record.
  • A sponsor wants the finding to remain reviewable after publication.
  • Subject response, corrections, and later revalidation must be visible.

The question

What does the recorded evidence establish about the specified public claim?

What we test

  • Specified claims and exclusions
  • Reproduction, replication, and method quality
  • Sensitivity and uncertainty
  • Subject response

What you receive

  • Versioned public report
  • Evidence labels
  • Reviewer statement
  • Correction and revalidation record

What to send

  • Specified public claims and exclusions
  • The completed reproduction, replication, or method-review record
  • Publication, factual-review, and confidentiality constraints

Publication terms, factual review, correction procedure, and sponsor disclosure are fixed before testing begins. A sponsor may fund the work but cannot control the method or conclusion.

For: Researchers, laboratories, journals, conferences, and public-interest sponsors

Discuss an independent report

For decisions

Test a technical premise before capital or adoption depends on it.

For funding, investment, acquisition, procurement, and R&D decisions.

Two technical evaluators test computing hardware with measurement equipment and execution logs.
Testing whether a research-derived system's claimed advantage survives controlled comparison.

Test the technical premise before you commit

Independent testing of the research claim or technical advantage behind an investment, acquisition, procurement, funding, or adoption decision.

See technical due diligence

The label can wait. Send the paper, artifact, or claim for a fit check.