Learn by running rnaseq-de
Build one useful,
verifiable skill.
Inspect rnaseq-de. Run it. Check it.
Then scope your own contribution.
Manuel Corpas · ClawBio
01 / The outcome
Make the result checkable
Define the input.
Know what the skill accepts.
Inspect the output.
Know what it produced.
Show the check.
Know why you trust it.
An agent can execute a command. You still own the question, the design and the judgement.
02 / Inspect
Start with the actual folder
skills/rnaseq-de/
SKILL.md
rnaseq_de.py
examples/
demo_counts.csv
demo_metadata.csvRead the contract in the instructions.
Inspect the data before running it.
Find the implementation that produces the result.
02 / Inspect
Two tables define this example
01 / Counts
10 genes × 6 samples
Genes in rows. Samples in columns.
Count-like values.
02 / Sample metadata
sample · condition · batch
Matching sample identifiers.
Control or treated, with recorded batch.
Counts + metadata → differential expression
Bundled toy data: 10 genes, 6 samples. No biological discovery is claimed.
02 / Inspect
Name the comparison before running
~ batch + condition
Contrast: condition,treated,control · Backend: pydeseq2
Positive log2 fold change: higher in treated.
Does the experimental design let you separate batch from condition?
03 / Run
Run the skill with a named backend
MPLBACKEND=Agg python skills/rnaseq-de/rnaseq_de.py \
--demo \
--backend pydeseq2 \
--output /tmp/clawbio-workshop-rnaseq
Run from the checkout root with a prepared environment.
Use a fresh output folder for each run.
Rehearsed with PyDESeq2 0.5.4. Setup, source pin and full commands
04 / Verify
Read the table before the plot
- Match samplesCounts and metadata agree.
- Confirm executionresult.json names pydeseq2.
- Check directionGeneA positive, GeneB negative in this fixture.
- Check identifiersAll 10 expected genes appear.
Directional sanity checks support this demonstration. They do not validate every dataset or model assumption.
04 / Verify
A plot still needs an explanation
Actual toy rehearsal output. Inspect the table and warnings before interpreting significance.
04 / Verify
A successful run has limits
manifest hashes verified
in the separate checksum file
Statistical warnings
Low residual degrees of freedom and numerical issues in the tiny fixture.
Incomplete provenance fieldsinput_checksum and datasets were empty in result.json.
Execution success is not statistical or biological validation.
05 / Challenge
Check when the skill should stop
Remove one sample row from a copy of the metadata.
Run with the unchanged count table.
Metadata missing samples
Restore the correct metadata or stop.
Do not invent the missing label.
This specific rejection was tested. It is not evidence that every invalid input is handled.
06 / Connect
Connect skills through a real contract
Required columns: gene, log2FoldChange,
and padj or pvalue.
A new visualisation does not change the inference.
Check identifiers, units and meaning as well as file format.
Optional demonstration. Pinned source
07 / Contribute
Find a gap before building
Inspect existing skills and tests. Reuse what already works.
- A small adapter with an equivalence check.
- A regression fixture for a reproducible failure.
- A missing validation check, if it is truly missing.
- A documentation fix or a precise bug report.
These are candidate contribution types, not claims of known gaps in rnaseq-de.
07 / Contribute
Write your smallest useful contribution
Define it
- Who needs the result?
- Exact input and output.
- Existing tool to reuse.
- Work you will leave out.
Make it testable
- One known-answer check.
- One invalid-input check.
- Builder and reviewer.
- First runnable checkpoint.
07 / Contribute
Ask a partner to explain your scope
Given [input], we produce [output].
We know it worked when [check].
We leave out [excluded work].
If your partner has to invent a missing detail, make the scope more concrete.
Your demo
Show the input.
The result. The check.
Explain one limitation.
A reproducible failure can be a useful contribution.