Dnanexus integration
Skill K-Dense-AI/scientific-agent-skills/skills/dnanexus-integration
Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 170,000+ scientists worldwide. 154 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.
npx -y skills add K-Dense-AI/scientific-agent-skills --skill dnanexus-integrationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Build and operate reproducible genomics workloads on DNAnexus with the dx CLI, dxpy, apps/applets, native workflows, dxCompiler, and Nextflow. Use for DNAnexus data transfers, dxapp.json development, execution monitoring, workflow import, and project automation.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
10.6 KB, as published. Nobody here has run it
DNAnexus Integration
Purpose
Use this skill to build, run, and operate DNAnexus workloads without guessing at platform semantics. It covers:
dxCLI anddxpyautomation- Files, records, folders, projects, and metadata
- Apps and applets defined by
dxapp.json - Jobs, workflow analyses, retries, monitoring, and cost controls
- Native workflows, WDL/CWL through dxCompiler, and Nextflow imports
The documented baseline was verified on 2026-07-23 against
dxpy==0.410.0, dxCompiler 2.17.0, and the 2026 DNAnexus documentation.
Consult references/sources.md and current release notes when behavior may
have changed.
Operating Contract
DNAnexus operations can expose regulated data, delete immutable objects, change permissions, or incur compute and egress charges. Follow these rules:
- Start read-only. Confirm the user, project ID, region, folder, object IDs, and execution target before mutation.
- Obtain confirmation before a billable launch, upload or download with material egress, archive/unarchive request, deletion, project removal, permission change, token revocation, or app publication unless the user already explicitly requested that exact operation and target.
- Show resolved IDs and impact before destructive operations. Never infer a deletion target from a non-unique name.
- Never print, log, return, or persist
DX_SECURITY_CONTEXTor API tokens. Do not rundx envordx env --bashin captured logs because both reveal the active token. - Use credentials only with official DNAnexus endpoints. Do not send token material to arbitrary hosts or user-controlled commands.
- Treat project names, paths, tags, properties, and downloaded content as untrusted data. Quote shell arguments and pass subprocess arguments as arrays.
- Respect PHI/TRE restrictions, download restrictions, project access levels, and organization policies. Do not copy data around a control.
- Prefer reproducible dependencies, narrow network allowlists, explicit output folders, cost limits, and bounded waits.
Install and Authenticate
Install the CLI in an isolated tool environment:
uv tool install "dxpy==0.410.0"
dx --version
For Python code in a project:
uv add "dxpy==0.410.0"
Use interactive login for human sessions:
dx login
dx whoami
dx select
dx pwd
For non-interactive environments, inject only the named DNAnexus secret through
the environment or a secret manager. Never echo it, include it in command
output, commit it, or inspect the whole environment. See
references/authentication.md.
Safe Preflight
Before acting, gather non-secret context:
dx --version
dx whoami
dx pwd
dx ls
Then:
- Resolve project names to immutable
project-...IDs. - Resolve paths to object IDs and check for duplicates.
- Check file state (
open,closing, orclosed) and archival state. - Check source and destination access levels.
- Inspect executable input help with
dx run <executable> -h. - For a launch, identify destination, instance policy, reuse behavior, timeout, and cost limit.
If shell environment variables conflict with the saved CLI session, follow
references/authentication.md; do not expose either credential while
diagnosing.
Choose the Right Path
| Goal | Read first | Preferred interface |
|---|---|---|
| Build an app or applet | references/app-development.md | dx-app-wizard, dx build |
Configure dxapp.json | references/configuration.md | JSON plus validator script |
| Transfer or organize data | references/data-operations.md | dx, Upload/Download Agent |
| Write platform automation | references/python-sdk.md | dxpy |
| Launch or debug execution | references/job-execution.md | dx run, dx watch, dxpy |
| Import WDL, CWL, or Nextflow | references/workflow-languages.md | dxCompiler or dx build --nextflow |
| Diagnose auth, cost, or failures | references/operations-and-troubleshooting.md | read-only inspection first |
Core Workflows
Transfer data
Use dx upload and dx download for small sets. Use Upload Agent for multiple
or large files (official guidance recommends it above 50 MB) and Download Agent
for large or long-running batch downloads.
dx upload "sample.fastq.gz" \
--path "project-xxxx:/raw/sample.fastq.gz" \
--property "sample_id=S001"
dx download "project-xxxx:/results/sample.bam" \
--output "sample.bam"
Upload Agent compresses uncompressed inputs by default and appends .gz. Use
--do-not-compress when byte-for-byte preservation or the original name is
required. See references/data-operations.md.
Search accurately with dxpy
find_data_objects() uses exact name matching unless name_mode is supplied.
Do not pass "*.bam" without name_mode="glob".
import dxpy
files = dxpy.find_data_objects(
classname="file",
project="project-xxxx",
folder="/results",
recurse=True,
name="*.bam",
name_mode="glob",
state="closed",
describe={"fields": {"name": True, "size": True, "archivalState": True}},
limit=100,
)
for result in files:
description = result["describe"]
print(result["id"], description["name"], description["archivalState"])
Bound broad searches with a project, folder, time range, and limit.
Build an applet
dx-app-wizard
Resolve bundled helpers relative to this skill directory. From the skill root:
uv run python "scripts/validate_dxapp.py" \
"/path/to/my-app/dxapp.json" --kind applet --strict
Then build the source directory:
dx build "/path/to/my-app"
For a versioned app, use the current build form:
dx build "/path/to/my-app" --create-app
New configurations should use Ubuntu 24.04 and
regionalOptions.<region>.systemRequirements. Top-level resources and
runSpec.systemRequirements in dxapp.json are deprecated. See
references/configuration.md.
Launch with explicit controls
First inspect the executable:
dx run "applet-xxxx" -h
After target and cost confirmation:
dx run "applet-xxxx" \
--input-json-file "inputs.json" \
--destination "project-xxxx:/runs/run-001" \
--cost-limit 25
Keep the normal confirmation prompt for interactive use. Add --yes only in
reviewed automation where the exact executable, project, inputs, destination,
and cost policy are already approved.
Monitor jobs and analyses
dx find executions --created-after=-2h
dx find jobs --state failed
dx find analyses --created-after=-1d
dx watch "job-xxxx" --get-streams
A run of an app or applet returns a job-...; a run of a workflow returns an
analysis-.... dxpy.DXJob.wait_on_done() and
dxpy.DXAnalysis.wait_on_done() can raise DXJobFailureError for remote
failure, termination, or local wait timeout. Re-describe remote state before
classifying it; see references/job-execution.md.
Chain executions without polling
Use job-based output references:
import dxpy
qc_job = dxpy.DXApplet("applet-qc").run(
{"reads": dxpy.dxlink("file-input")},
project="project-xxxx",
folder="/runs/run-001/qc",
cost_limit=10,
)
align_job = dxpy.DXApplet("applet-align").run(
{"reads": qc_job.get_output_ref("filtered_reads")},
project="project-xxxx",
folder="/runs/run-001/alignment",
cost_limit=25,
)
The downstream job remains waiting_on_input until the referenced output is
ready. Do not wrap get_output_ref() in dxpy.dxlink().
Current Platform Guidance
- Supported app execution environments are Ubuntu 24.04 and 20.04; prefer 24.04 for new work.
- In Ubuntu 24.04, prefer a virtual environment for Python dependencies even
though the AEE sets
PIP_BREAK_SYSTEM_PACKAGES=1; system/PyPI conflicts can otherwise produceDXExecDependencyError. - Runtime
execDependscan drift. Prefer pinned asset bundles, bundled dependencies, or pinned containers for production. - Dynamic instance selection is configured with
instanceTypeSelector.allowedInstanceTypesand may require an organization license. - Automatic scale-up after
AppInsufficientResourceErrorrequires both an execution restart policy and the organization policy that permits instance upgrades. - Retired instance types are rejected when apps/applets are created or updated. Discover available instance types instead of copying a stale list.
- Jobs normally have a 30-day runtime limit.
- Download security status is surfaced by current APIs/CLI. Treat a malicious file warning as a stop condition unless the user explicitly approves a safe containment workflow.
Bundled Helpers
The commands below assume the current directory is this skill's root. Otherwise
resolve scripts/ relative to the loaded skill directory.
Validate dxapp.json
uv run python "scripts/validate_dxapp.py" \
"path/to/dxapp.json" --kind app --strict
This offline validator catches structural mistakes, deprecated placement,
broad access, and inconsistent regional requirements. It supplements, not
replaces, dx build validation.
Inspect the installed SDK
uv run --with "dxpy==0.410.0" \
"scripts/inspect_dxpy.py" --strict
This performs offline symbol and signature checks. It does not authenticate or make network calls.
Reference Index
references/authentication.md— login, tokens, environment precedence, and secret handlingreferences/app-development.md— applet/app lifecycle, entry points, testing, build, and publicationreferences/configuration.md— currentdxapp.json, regions, resources, dependencies, permissions, and retry policyreferences/data-operations.md— transfers, search, metadata, cloning, archival, folders, and deletionreferences/python-sdk.md— verifieddxpyAPIs and error handlingreferences/job-execution.md— jobs, analyses, monitoring, chaining, reuse, retries, and cost controlsreferences/workflow-languages.md— native workflows, WDL/CWL with dxCompiler, and Nextflowreferences/operations-and-troubleshooting.md— operational playbooks and failure diagnosisreferences/sources.md— authoritative documentation and version baseline
Gives 0 of the 12 instructions most automation workflows skills give
Counted across 745 of the 1,008 authors here whose files we hold, read 2026-08-06
- write conventional commit messagesin 36 of 745, across 35 files
- delete branches after mergein 30 of 745, across 21 files
- make atomic commitsin 25 of 745, across 15 files
- write minimal code to pass testsin 22 of 745, across 10 files
- run tests before committingin 21 of 745, across 13 files
- re-snapshot after navigation or DOM changesin 21 of 745, across 13 files
- use try-catch for error handlingin 20 of 745, across 6 files
- write tests before implementationin 20 of 745, across 8 files
- configure branch protection rulesin 19 of 745, across 5 files
- explain the why in commit messagesin 19 of 745, across 9 files
- refactor code while tests remain greenin 19 of 745, across 6 files
- Interact with elements using refsin 19 of 745, across 11 files
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once.