Sedona use-case notebooks¶
The notebooks at the root of this directory (docs/usecases/*.ipynb) are bundled into the Sedona docker image at /opt/workspace/examples/ and rendered on the docs site via mkdocs-jupyter. The legacy/ subdirectory holds older notebooks kept for backward-compatible URL access only β they are not shipped in the image. The contrib/ subdirectory holds community-submitted examples that are also not shipped.
Contract for shipped notebooks¶
Every notebook at the root of this directory must:
- Run end-to-end in the docker image with default resources (
DRIVER_MEM=4g,EXECUTOR_MEM=4g,local[*]). No external clusters, no GPUs. - Finish in well under the 900-second per-notebook timeout enforced by
docker/test-notebooks.sh. Aim for under 2 minutes wall-clock on a laptop with a warm dataset cache. - Use only data that ships in the image (
docs/usecases/data/) or is fetched from a public, anonymous-readable URL. No credentials, no private buckets. - Tag network-dependent notebooks so they can be skipped in sandboxed CI:
<!-- requires-network: true -->
docker/test-notebooks.sh greps for that exact string and skips the notebook when SEDONA_NOTEBOOK_OFFLINE=1 is set in the environment.
5. Use data/... relative paths for shipped data and .master("spark://localhost:7077") for the SedonaContext β the test harness rewrites both into the form needed for local[*] test mode (examples/data/..., local[*]).
How CI verifies the contract¶
.github/workflows/... builds the docker image and runs:
docker run --rm sedona:dev /opt/sedona/docker/test-notebooks.sh
That script:
- iterates every
*.ipynbat the root of/opt/workspace/examples/; - for each, runs
jupyter nbconvert --to python --stdout, applies sed rewrites (spark://...βlocal[*],data/...βexamples/data/...), strips IPython magics, then runs the resulting.pyunderpython3with a 900s timeout andset -o pipefailso a notebook crash or timeout cannot be silently misreported as a pass; - skips notebooks tagged
requires-network: truewhenSEDONA_NOTEBOOK_OFFLINE=1; - exits non-zero on any failure.
How to verify a notebook locally before opening a PR¶
The fast feedback loop is the docker build:
# from the repo root
docker build -f docker/sedona-docker.dockerfile -t sedona:dev .
docker run --rm sedona:dev /opt/sedona/docker/test-notebooks.sh
docker run --rm -e SEDONA_NOTEBOOK_OFFLINE=1 sedona:dev /opt/sedona/docker/test-notebooks.sh
Both invocations must exit 0. The second variant proves your notebook either runs offline or is correctly tagged requires-network: true.
For faster iteration during notebook authoring, you can replicate the harness without a docker rebuild by running the notebook directly in a Python environment that matches the image's runtime:
- Python 3.10+ (image uses 3.12)
pyspark==4.0.1apache-sedona==1.9.1keplergl==0.3.7,geopandas,pyarrow,pandas,matplotlib,fiona==1.10.1JAVA_HOMEpointing at a JDK 17 install
Then export the Sedona Maven coordinates (the docker image bakes the JAR in directly; locally you let Maven pull it):
export PYSPARK_SUBMIT_ARGS="--packages org.apache.sedona:sedona-spark-shaded-4.0_2.13:1.9.1,org.datasyslab:geotools-wrapper:1.9.1-33.5 --driver-memory 4g pyspark-shell"
Convert and run:
jupyter nbconvert --to python docs/usecases/<your-notebook>.ipynb --stdout > /tmp/run.py
sed -i '' 's|\.master("spark://[^"]*")|.master("local[*]")|g' /tmp/run.py
( cd docs/usecases && python /tmp/run.py )
Successful runs match what docker/test-notebooks.sh will see in CI.
Adding a new shipped notebook¶
- Drop the
.ipynbat the root of this directory. - If it needs new shipped data, add it under
data/and verify total added bytes stay under ~25 MB unless there's a strong reason. - Verify it passes the contract above (network on AND off paths).
- Open a PR with
[GH-####]referencing the relevant issue.