Testing

Unit testing is one of the single most important ingredients in writing maintainable code. Once you trust your tests, you will be set free to pull things apart, put them back together again, and have confidence that you have minimized the risk of breaking your downstream customers with your next update.

For this reason, TTS strives to have 100% coverage on all of our libraries, and has included test coverage reports in our unit test matrix (see https://github.com/NASA-JPL-Teamtools-Studio/tts_ci_cd/)

TTS Testing Conventions

  • Use pytest. The world won't end if you're contributing something using unittest, but in general, that is our style.
  • Unless there is a compelling reason not to, try to match your test architecture one-to-one with the architecture with your code architecture.
    • One test file per code file
    • One test per function/method
    • When using supporting files, structure your test file directory so it is very easy to track which file belongs to which test
    • When using supportnig files, only share files between tests if they are extremely similar
  • Tests written by GenAI are better than no tests at all.
    • GenAI is a powerful tool in writing tests, but shoulnd't be used naively
    • If you write a test with GenAI, you should mark it as unreviewed until you have clearly walked through and understand it (https://github.com/NASA-JPL-Teamtools-Studio/tts_ci_cd/issues/1)
      • Note: we have not followed this convention to date, but will soon start.
      • To begin, all tests should be marked as unreviewed even though many were written by humans
  • Know how code coverage actually works
    • pytest-cov gives you credit for exerising a line of code
    • It has no way to understand that you have written a valid test
    • You could exercise every line and just assert True at the end and get 100% coverage
    • See https://github.com/NASA-JPL-Teamtools-Studio/tts_ci_cd/issues/2 for more information

Test Matrix

Although it is impossible to test in every permutation of Python version and dependency version that out customers may be using, TTS understands that teams using our code may be using different configurations than we are. "It works on my machine" is not sufficient.

For this reason, we have developed a matrix testing strategy as follows: * Run tests against every major version of Python. * In general this is the most recent minor version of each major version * The exception is 3.6.8, which we use as our oldest supported version * 3.6.8 still comes as default on many systems * Allow pip to resolve whichever lastest version of each library is available for each test * Use a clean Docker container for testing each library. Clone it and install all of its dependencies at run time. * These libraries are best used as a whole suite, but are designed to be used independently of each other, so we smoke out any implicit dependencies here so we can fix them * Use unit testing as a time to run other housekeeping tools * Security flaws via pip-audit and bandit * Test coverage via pytest-cov * Documentation coverage via a home-grown script * Check if every function, method, class, and arg/kwarg has a docstring

See https://https://nasa-jpl-teamtools-studio.github.io/tts_ci_cd/ for latest matrix test results

Inspection Tests and the @pytest.mark.inspection Marker

Some tests produce visual output (HTML tables, styled reports, CSV exports) that cannot be fully verified by machine assertion. These are inspection tests and they carry a distinct marker:

@pytest.mark.inspection
class TestMyFeatureInspection:
    def test_write_inspection_html(self):
        ...

Rules for inspection tests: - Mark every inspection test class (or individual test) with @pytest.mark.inspection. This is the signal to both CI and humans that the test requires eyes, not just an assertion. - The test generates an artifact to src/<package>/test/core/test_files/<feature>_inspection.html and calls check_inspection_hash(path) at the end. - check_inspection_hash fails until a human runs python src/<package>/test/certify.py (no args to see status, --certify to stamp approval) and commits the resulting .sha256 sidecar file. - Never update .sha256 files from CI or AI agents. They are the human stamp of approval.

Entry point for reviewers: run python src/<package>/test/certify.py to generate test_files/inspection_status.html — a self-contained dashboard listing every artifact with its review state and a clickable link. Point reviewers here, not at the test output.

See docs/adr/001-open-source-over-saas.md for why we built this rather than adopting a SaaS visual regression tool.

AI-Generated Tests and the @pytest.mark.unreviewed_ai Marker

AI agents and tools can generate tests quickly, but a passing test is not the same as a correct test. unreviewed_ai keeps the team honest by making it possible to compare coverage between tests a human has reviewed and tests a human has not.

@pytest.mark.unreviewed_ai
class TestMyFeatureGenAI:
    def test_some_behavior(self):
        ...

Rules: - Mark any test written or substantially generated by AI with @pytest.mark.unreviewed_ai until a human has read every line and is confident it tests the right thing. - Remove the marker once a human has reviewed the test — that is the signal it has been promoted to a trusted test. - All existing tests that pre-date this convention should initially carry @pytest.mark.unreviewed_ai to establish a baseline. See tts_ci_cd #1 for migration status.

Why this matters: pytest-cov gives credit for exercising a line — it cannot tell you whether the assertion is meaningful. AI-generated tests can achieve 100% line coverage while asserting nothing useful. Tracking unreviewed_ai coverage separately surfaces this risk.

CI Job Structure

Three separate jobs are required:

Job Command Blocks merge?
Reviewed unit tests pytest -m "not inspection and not unreviewed_ai" Yes
AI-generated tests pytest -m unreviewed_ai Yes
Inspection tests pytest -m inspection Yes

Running them separately makes it immediately clear in PR status which category of failure you are dealing with.

Registering both markers: add to pyproject.toml:

[tool.pytest.ini_options]
markers = [
    "inspection: test produces visual output requiring human inspection and sha256 certification before passing",
    "unreviewed_ai: test was written or generated by AI and has not yet been reviewed line-by-line by a human",
]

Edit/Comment on GitHub