Skip to content

Repository files navigation

seisfetch

Cloud-first seismic waveform access for EarthScope, SCEDC, NCEDC, GeoNet, and fallback FDSN services.

seisfetch is built around one core path:

cloud archive / HTTP  ->  raw miniSEED bytes  ->  pymseed  ->  numpy arrays
                                                      |
                                                      +-> xarray
                                                      +-> zarr
                                                      +-> ObsPy
                                                      +-> Earth2Studio adapters

The design goal is simple:

  • prefer cloud-native archive access first
  • use FDSN second, as a fallback path
  • decode miniSEED straight into numpy without requiring ObsPy in the main pipeline
  • stay compatible with sparse sensor and Earth2Studio workflows

Overall Structure

The intended acquisition order is:

  1. s3_open for SCEDC, NCEDC and GeoNet open buckets
  2. s3_auth for EarthScope S3 access
  3. fdsn only when the archive-backed path is unavailable or the network is not served from those buckets

At the package level, the main interfaces are:

  • SeisfetchClient.get_raw() -> raw miniSEED bytes
  • SeisfetchClient.get_numpy() -> TraceBundle of numpy arrays
  • SeisfetchClient.get_xarray() -> xarray.Dataset
  • SeisfetchClient.get_waveforms() -> ObsPy Stream
  • SeismicDataFrameSource / SeismicDataSource -> Earth2Studio adapters over data you already fetched
  • SeisfetchLiveSource -> time-indexed Earth2Studio DataSource that fetches on demand

Workflow

1. Cloud-first archive access

seisfetch routes by network code:

Network family Preferred source Backend
CI, other SCEDC-routed networks SCEDC open S3 s3_open
BK, other NCEDC-routed networks NCEDC open S3 s3_open
NZ GeoNet open S3 s3_open
IU, UW, TA, other EarthScope-routed networks EarthScope S3 s3_auth

Networks served by both SCEDC and NCEDC (NC, NP, ...) route to NCEDC.

Archive details:

Archive Bucket Region Auth
EarthScope earthscope-geophysical-data us-east-2 EarthScope SDK credentials
SCEDC scedc-pds us-west-2 none
NCEDC ncedc-pds us-west-2 none
GeoNet geonet-open-data ap-southeast-2 none

Notes:

  • SCEDC, NCEDC and GeoNet are per-channel archives, so you should pass channel=....
  • GeoNet channels always carry a numeric location code (10, 20, ...), so location= is required there — a blank location raises rather than silently missing data.
  • EarthScope stores station-day miniSEED objects and currently requires authenticated access through earthscope-sdk.

2. FDSN second

Use backend="fdsn" when:

  • the desired network is not available from EarthScope / SCEDC / NCEDC / GeoNet S3
  • the archive-backed attempt fails and you want a fallback provider
  • you need a non-US provider such as GEOFON, INGV, ETH, ORFEUS, etc.

This keeps the default workflow archive-first instead of HTTP-first.

3. Decode directly to numpy

The central decode path is:

  • fetch raw miniSEED bytes
  • decode with pymseed
  • work in numpy immediately

ObsPy is not required for this path.

4. Convert only when needed

Once you have a TraceBundle, you can convert to:

  • xarray.Dataset for labeled arrays and ML/data workflows
  • zarr for chunked cloud/local persistence
  • ObsPy Stream for classical seismology tooling
  • Earth2Studio adapters for sparse sensor and digital twin/data assimilation workflows

Why This Package

Existing seismic client workflows often couple:

  • transport
  • decode
  • metadata
  • downstream object model

seisfetch separates those concerns:

  • transport: S3 or HTTP
  • decode: miniSEED -> numpy
  • output: xarray / zarr / ObsPy / Earth2Studio only when requested

That makes it useful for:

  • cloud-native waveform mining
  • sparse sensor ingestion
  • foundation-model training pipelines
  • Earth2Studio interoperability
  • data assimilation and digital twin workflows

ObsPy-Free, and What That Buys

ObsPy is an optional extra here, not a dependency of the data path. The evaluation behind that choice — including whether it changes the science — is written up in docs/noisepy-obspy-replacement-report.md, with every number traceable to a committed JSON under benchmarks/results/ or a test in tests/precision/.

seisfetch obspy stack
Parse an 11 MB Steim2 channel-day 21.4 ms 37.2 ms
Cold import 0.08 s 0.13 s
Parse peak memory 27.7 MB 52.0 MB
Installed footprint 80.4 MB 311.4 MB
arm64 Linux install wheels needs gcc

The footprint line is the cloud argument: AWS Lambda caps a layer at 250 MB, so the ObsPy stack does not fit and seisfetch does. ObsPy publishes no linux/aarch64 wheels, so on Graviton it compiles from source; seisfetch and pymseed install from wheels.

The science does not change

Identical archive bytes were pushed through (A) obspy.read plus NoisePy's own preprocess_raw and (B) seisfetch's parser plus the numpy/scipy ports in seisfetch.contrib.noisepy_adapter, then through NoisePy's own compute_fft and correlate. The pass criterion is bit-identity, not closeness:

Harness Result
Single-station CI.PASC, EN/EZ/NZ/ZZ at 40 sps max abs diff 0.0
Cross-station SCEDC x NCEDC x EarthScope at 20 sps max abs diff 0.0
dv/v stretching same grid cell on all pairs

The cross-station run matters because it exercises the Fourier-resample and sub-sample-alignment branches inside the chain. Tables, plots and the harnesses: benchmarks/RESULTS.md.

Scope note: this bit-identity result covers the preprocessing chain at rm_resp=NO. Response removal is validated separately, below.

Response removal without evalresp

Response removal was the one ObsPy capability the NoisePy migration still needed. seisfetch.contrib.response provides it in ~577 lines of numpy and stdlib xml.etree — no evalresp C library, no ObsPy, no lxml, and no scipy in that module. ObsPy has no pure-Python response evaluator (remove_response calls compiled evalresp), so every stage was re-derived and then checked against the compiled implementation:

Check Result
evaluate_response(mode="full") vs compiled evalresp — both CI.PASC epochs, VEL/ACC/DISP, 1 mHz–19.9 Hz max rel diff 1.6e-10
remove_response_np vs Trace.remove_response — real 6.9M-sample Tohoku day, water_level=60, pre_filt 6.6e-16 of peak
Same, CI.PASC demo hour in notebook 06 7.6e-16 of peak
Deconvolve that Tohoku day 1.9 s vs ObsPy 3.6 s

Two evaluation modes: mode="full" is evalresp-equivalent (all stages, analog and digital poles/zeros, FIR/Coefficients with DC normalization and the Decimation/CorrectionApplied phase advance). mode="paz" is the SACPZ shortcut — 0.7–1.3 % error below 4 Hz but ~23 % by 16 Hz, since the FIR anti-alias roll-off is unmodeled; don't use it above ~Nyquist/3.

Two deconvolution styles: remove_response_np ports ObsPy's water-level method for drop-in equivalence, and translate_resp_np follows SeisIO.jl's translation approach, taking stabilization from the target response's own roll-off instead of a water level.

Defective metadata fails loudly — zero or missing gains, degenerate normalization references, zero-sum FIR stages, polynomial and ResponseList stages all raise with the stage number named, never a silent NaN or unity gain. Not implemented, and raising rather than approximating: IIR Coefficients stages with denominators, polynomial (blockette-62) responses, ResponseList stages. Metadata is StationXML only; RESP and SACPZ files are not parsed.

Full derivation, the conditional-A0 finding about evalresp's normalization rule, and the SeisIO comparison: docs/response-removal-design.md. Tutorial: notebooks/05_response_removal.ipynb.

Quick Start

Open S3: SCEDC / NCEDC

from seisfetch import SeisfetchClient

client = SeisfetchClient(backend="s3_open")

bundle = client.get_numpy(
    "CI",
    "ABL",
    channel="BHZ",
    starttime="2024-01-15T00:00:00",
    endtime="2024-01-15T00:10:00",
)

print(bundle.ids)
arrays = bundle.to_dict()

Open S3: GeoNet (New Zealand)

NZ auto-routes to the GeoNet open-data bucket. GeoNet channels carry a numeric location code, so pass location=:

bundle = SeisfetchClient(backend="s3_open").get_numpy(
    "NZ",
    "WEL",
    location="10",
    channel="HHZ",
    starttime="2022-01-02T00:00:00",
    endtime="2022-01-02T00:10:00",
)

EarthScope S3 (authenticated)

Requires earthscope-sdk and an EarthScope account that has been granted the s3-miniseed role. See EarthScope Credentials.

client = SeisfetchClient(backend="s3_auth")

bundle = client.get_numpy(
    "IU",
    "ANMO",
    location="00",
    channel="BHZ",
    starttime="2024-01-15T00:00:00",
    endtime="2024-01-15T00:01:00",
)

EarthScope via FDSN (no S3 role required)

Works for any logged-in EarthScope account, including those without direct S3 access. Good for TA, IU, US, UW, _GSN, etc.:

client = SeisfetchClient(backend="fdsn", providers="EARTHSCOPE")

bundle = client.get_numpy(
    "TA",
    "034A",
    channel="BHZ",
    starttime="2010-06-01T00:00:00",
    endtime="2010-06-01T00:05:00",
)
print(bundle.ids)            # ['TA.034A..BHZ']
print(bundle.to_dict()[bundle.ids[0]].shape)

Station discovery

get_stations() queries fdsnws-station and auto-routes to SCEDC, NCEDC, or EarthScope based on the network code. No ObsPy required.

rows = client.get_stations(
    "TA", channel="BHZ",
    starttime="2010-06-01", endtime="2010-06-02",
)
for r in rows[:3]:
    print(r["Network"], r["Station"], r["Channel"], r["Latitude"], r["Longitude"])

FDSN fallback

client = SeisfetchClient(backend="fdsn", providers="GEOFON")

bundle = client.get_numpy(
    "GE",
    "BKB",
    channel="BHZ",
    starttime="2011-03-11T06:00:00",
    endtime="2011-03-11T06:05:00",
)

xarray output

ds = client.get_xarray(
    "CI",
    "ABL",
    channel="BHZ",
    starttime="2024-01-15T00:00:00",
    endtime="2024-01-15T00:10:00",
)

Earth2Studio-compatible sparse sensor output

from datetime import datetime
from seisfetch import SeismicDataFrameSource

df_source = SeismicDataFrameSource(bundle)
df = df_source(datetime(2024, 1, 15), list(ds.data_vars))

Metadata table + zarr sidecar

from seisfetch import bundle_to_metadata_table, to_zarr, write_metadata_csv

metadata_table = bundle_to_metadata_table(bundle)
to_zarr(bundle, "quickstart.zarr", metadata=metadata_table)
write_metadata_csv(metadata_table, "quickstart.zarr")

Dependency Tuning

You do not need every dependency for every workflow.

Minimal core

Use this when you only want miniSEED -> numpy from S3:

pip install seisfetch

Includes:

  • numpy
  • boto3
  • pymseed

Add FDSN fallback

Use this if you want direct HTTP fallback providers:

pip install "seisfetch[fdsn]"

Adds:

  • httpx

Add authenticated EarthScope S3

Use this if you need EarthScope archive access:

pip install "seisfetch[auth]"
pip install earthscope-cli

Adds:

  • earthscope-sdk
  • EarthScope CLI login flow

Add xarray

Use this for labeled arrays and ML pipelines:

pip install "seisfetch[xarray]"

Add metadata tables

Use this if you want canonical metadata tables, CSV sidecars, or metadata-aware zarr output:

pip install "seisfetch[pandas]"

Add zarr

Use this for chunked persistent stores:

pip install "seisfetch[zarr]"

Add ObsPy

Use this only if you need ObsPy interop or ObsPy-backed FDSN behavior:

pip install "seisfetch[obspy]"

Suggested dependency bundles

Need Install
S3 open data -> numpy only pip install seisfetch
Metadata table / metadata.csv export pip install "seisfetch[pandas]"
Archive-first + FDSN fallback pip install "seisfetch[fdsn]"
EarthScope + SCEDC + NCEDC pip install "seisfetch[auth]" and pip install earthscope-cli
xarray / ML / Earth2Studio-style workflows pip install "seisfetch[xarray]"
zarr persistence pip install "seisfetch[zarr]"
ObsPy interop pip install "seisfetch[obspy]"
Most common research stack pip install "seisfetch[fdsn,auth,xarray,zarr,obspy]"

Installation Modes

pip

For scripts and lightweight pipelines:

pip install seisfetch

From source:

git clone https://github.com/Denolle-Lab/seisfetch
cd seisfetch
pip install .

pixi

For notebook work and development:

git clone https://github.com/Denolle-Lab/seisfetch
cd seisfetch
pixi install
pixi install -e notebooks
pixi run -e notebooks kernel-install

The notebook environment is the intended Jupyter environment for this repo.

EarthScope Credentials

EarthScope has two access tiers:

  1. FDSN web service (backend="fdsn", providers="EARTHSCOPE") — available to any logged-in account. No special role required.
  2. Direct S3 (backend="s3_auth") — requires the s3-miniseed IAM role to be granted on your account. This is faster and cheaper when you are running in us-east-2.

Setup

pip install "seisfetch[auth]"
pip install earthscope-cli
es login

The [auth] extra already installs earthscope-sdk, so there is no separate SDK install step.

Verify (CLI)

# 1. Confirm you are logged in
es user get-profile                # prints your name, email, institution

# 2. Confirm direct-S3 role is granted (only needed for backend="s3_auth")
es user get-aws-credentials        # prints temporary AWS keys, or an error

A response of "You are not allowed to assume role 's3-miniseed'" or UnauthorizedError means your account is logged in but direct S3 is not enabled yet. Use backend="fdsn", providers="EARTHSCOPE" in the meantime and email data-help@earthscope.org to request the s3-miniseed role.

Verify (Python)

from earthscope_sdk import EarthScopeClient

with EarthScopeClient() as client:
    print(client.user.get_profile())
    try:
        creds = client.user.get_aws_credentials()
        print("S3 role granted:", creds.aws_access_key_id[:8])
    except Exception as exc:
        print("S3 role NOT granted:", exc)

Headless / CI

es user get-refresh-token
export ES_OAUTH2__REFRESH_TOKEN="<your-refresh-token>"

Architecture

SeisfetchClient
|
+- get_raw()        -> raw miniSEED bytes
+- get_numpy()      -> TraceBundle (numpy)
+- get_xarray()     -> xarray.Dataset
+- get_waveforms()  -> ObsPy Stream
|
+- backend="s3_open"
|  +- SCEDC open bucket
|  +- NCEDC open bucket
|  +- GeoNet open bucket
|  +- auto-routing by network code
|
+- backend="s3_auth"
|  +- EarthScope S3 via earthscope-sdk credentials
|
+- backend="fdsn"
|  +- single-provider HTTP client
|  +- multi-provider fan-out client
|
+- backend="obspy_fdsn"
   +- ObsPy-backed fallback for harder FDSN cases

Earth2Studio Compatibility

The package includes adapters in seisfetch.earth2 for Earth2Studio-style usage:

  • SeismicDataSource — wraps a bundle you already fetched
  • SeismicDataFrameSource — sparse sensor table; auto_coords=True fills station lat/lon from the FDSN station service
  • SeisfetchLiveSource — time-indexed DataSource that fetches on demand
  • bundle_to_earth2

These are intended for:

  • sparse sensor tables
  • observation pipelines
  • foundation-model data preparation
  • Earth2Studio / digital twin workflows

Typical path:

miniSEED -> numpy -> xarray / sparse dataframe -> Earth2Studio adapter

Live source

SeisfetchLiveSource has the shape every other Earth2Studio source (GFS, ERA5, ...) has: you call it with timestamps and it fetches, auto-routing per network across all four archives and caching day bundles in memory.

from datetime import datetime
from seisfetch.earth2 import SeisfetchLiveSource

source = SeisfetchLiveSource(
    channels=["CI.PASC..BHZ", "BK.PKD.00.BHZ", "II.PFO.00.BHZ"],
    window_s=3600,
    calibrate="gain",
)
da = source(datetime(2022, 1, 2, 6))   # -> (time, variable, sample) DataArray

Channels are NET.STA.LOC.CHA strings; the returned variable coordinate spells them with underscores (CI_PASC__BHZ), which is also what the optional variable= argument accepts. All channels in one call must share a sampling rate — request mixed rates (a 40 sps BHZ alongside a 100 sps HHZ) in separate calls.

Physical units are required — this source never returns raw counts:

  • calibrate="gain" (default): divide by the channel's total sensitivity from the FDSN station service. Exact at the reference frequency, one metadata request per channel, no extra dependencies.
  • calibrate="response": full spectral deconvolution through seisfetch.contrib.response (StationXML fetch plus an evalresp-equivalent evaluator). Agrees with ObsPy to 7.6e-16 of peak amplitude on real data.

fetch() is genuinely async (asyncio.to_thread), so pipelines can overlap this source with others.

Recipes

A few common end-to-end tasks the package is designed for.

Mine a single station-day to numpy

from seisfetch import SeisfetchClient

bundle = SeisfetchClient(backend="s3_open").get_numpy(
    "CI", "ABL", channel="BHZ",
    starttime="2024-01-15T00:00:00",
    endtime="2024-01-16T00:00:00",
)
data = bundle.to_dict()["CI.ABL..BHZ"]   # int32 numpy array, full day

Bulk fetch many station-channels in parallel

requests = [
    {"network": "CI", "station": s, "channel": c,
     "starttime": "2025-04-14T17:07:30", "endtime": "2025-04-14T17:10:30"}
    for s in ["ABL", "SDD", "PASC"]
    for c in ["BHZ", "BHN", "BHE"]
]
client = SeisfetchClient(backend="s3_open")
summary = client.get_numpy_bulk(requests, max_workers=8)
print(summary.succeeded, "/", summary.total)

Missing station-channel combinations are reported in summary.failed and do not abort the run.

Archive-first with FDSN fallback

from seisfetch import SeisfetchClient, route_network

def get_archive_first(net, sta, *, starttime, endtime, location="*", channel="*",
                       fallback="GEOFON"):
    dc = route_network(net)
    primary = "s3_open" if dc in {"scedc", "ncedc"} else "s3_auth"
    try:
        return SeisfetchClient(backend=primary).get_numpy(
            net, sta, location=location, channel=channel,
            starttime=starttime, endtime=endtime,
        )
    except Exception:
        return SeisfetchClient(backend="fdsn", providers=fallback).get_numpy(
            net, sta, location=location, channel=channel,
            starttime=starttime, endtime=endtime,
        )

Persist a multi-channel package to zarr with metadata sidecar

from seisfetch import bundle_to_metadata_table, to_zarr, write_metadata_csv

metadata = bundle_to_metadata_table(bundle)
to_zarr(bundle, "package.zarr", metadata=metadata)
write_metadata_csv(metadata, "package.zarr")

The resulting package.zarr/ contains channel groups plus a metadata/channel_table group readable with xarray.open_zarr(..., group="metadata/channel_table").

Earth2Studio sparse-sensor handoff

from datetime import datetime
from seisfetch import SeismicDataFrameSource, bundle_to_xarray

ds = bundle_to_xarray(bundle)
df = SeismicDataFrameSource(bundle)(datetime(2024, 1, 15), list(ds.data_vars))
df[["time", "variable", "network", "station", "channel",
     "sampling_rate", "amplitude_rms", "num_samples"]].head()

Command Line

Examples:

# SCEDC open S3 (no auth)
seisfetch download CI ABL -s 2024-01-15 -e 2024-01-15T01:00:00 -c BHZ -o data.mseed
seisfetch numpy    CI SDD -s 2024-06-01 -c BHZ -o data.npz
seisfetch zarr     CI ABL -s 2024-01-15 -c BHZ -o data.zarr

# EarthScope (requires `es login` and the s3-miniseed role)
seisfetch download IU ANMO -s 2024-01-15 -e 2024-01-15T01:00:00 -c BHZ --backend s3_auth -o anmo.mseed

# Routing and provider info
seisfetch info --route CI
seisfetch info --providers

# Bulk
seisfetch bulk requests.csv -o output/ -f npz

Notebooks

See notebooks/ for worked examples:

Notebook setup instructions are in notebooks/README.md.

Tests

pixi run test
pixi run test-cov
pixi run test-int

Dependencies

Package Status Role
numpy core array container
boto3 core S3 transport
pymseed core miniSEED decode
httpx optional [fdsn] FDSN HTTP client
pandas optional [pandas] canonical metadata tables and CSV export
xarray optional [xarray] labeled dataset output
zarr optional [zarr] persistent chunked storage
obspy optional [obspy] ObsPy interop and alternative FDSN backend
earthscope-sdk optional [auth] EarthScope S3 credentials

See THIRD_PARTY_NOTICES.md for attribution and licenses.

Benchmarks with plots: benchmarks/RESULTS.md (or the self-contained RESULTS.html). Changes: CHANGELOG.md. Response-removal tutorial: notebooks/05_response_removal.ipynb.

Citation

When using data accessed through seisfetch:

  • EarthScope: cite the network operators and NSF SAGE facility
  • SCEDC: doi:10.7909/C3WD3xH1
  • NCEDC: doi:10.7932/NCEDC
  • GeoNet: GNS Science, GeoNet open data (CC BY 4.0) — cite per GeoNet's data policy
  • other FDSN providers: cite the underlying network/provider

Software references:

License

MIT AND LGPL-3.0-only. The package is MIT (LICENSE) with one exception: seisfetch/contrib/obspy_ports.py contains numerically exact translations of ObsPy routines and is LGPL-3.0-only (ObsPy is © The ObsPy Development Team, LGPL v3). Using seisfetch as a library is unaffected; redistributors of modified versions of that one file must honor LGPL terms. Details in THIRD_PARTY_NOTICES.md.

About

Fast seismic miniSEED from EarthScope, SCEDC, NCEDC, and 37+ FDSN servers. No ObsPy required.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages