Cloud-first seismic waveform access for EarthScope, SCEDC, NCEDC, GeoNet, and fallback FDSN services.
seisfetch is built around one core path:
cloud archive / HTTP -> raw miniSEED bytes -> pymseed -> numpy arrays
|
+-> xarray
+-> zarr
+-> ObsPy
+-> Earth2Studio adapters
The design goal is simple:
- prefer cloud-native archive access first
- use FDSN second, as a fallback path
- decode miniSEED straight into numpy without requiring ObsPy in the main pipeline
- stay compatible with sparse sensor and Earth2Studio workflows
The intended acquisition order is:
s3_openfor SCEDC, NCEDC and GeoNet open bucketss3_authfor EarthScope S3 accessfdsnonly when the archive-backed path is unavailable or the network is not served from those buckets
At the package level, the main interfaces are:
SeisfetchClient.get_raw()-> raw miniSEED bytesSeisfetchClient.get_numpy()->TraceBundleof numpy arraysSeisfetchClient.get_xarray()->xarray.DatasetSeisfetchClient.get_waveforms()-> ObsPyStreamSeismicDataFrameSource/SeismicDataSource-> Earth2Studio adapters over data you already fetchedSeisfetchLiveSource-> time-indexed Earth2StudioDataSourcethat fetches on demand
seisfetch routes by network code:
| Network family | Preferred source | Backend |
|---|---|---|
CI, other SCEDC-routed networks |
SCEDC open S3 | s3_open |
BK, other NCEDC-routed networks |
NCEDC open S3 | s3_open |
NZ |
GeoNet open S3 | s3_open |
IU, UW, TA, other EarthScope-routed networks |
EarthScope S3 | s3_auth |
Networks served by both SCEDC and NCEDC (NC, NP, ...) route to NCEDC.
Archive details:
| Archive | Bucket | Region | Auth |
|---|---|---|---|
| EarthScope | earthscope-geophysical-data |
us-east-2 |
EarthScope SDK credentials |
| SCEDC | scedc-pds |
us-west-2 |
none |
| NCEDC | ncedc-pds |
us-west-2 |
none |
| GeoNet | geonet-open-data |
ap-southeast-2 |
none |
Notes:
- SCEDC, NCEDC and GeoNet are per-channel archives, so you should pass
channel=.... - GeoNet channels always carry a numeric location code (
10,20, ...), solocation=is required there — a blank location raises rather than silently missing data. - EarthScope stores station-day miniSEED objects and currently requires authenticated access through
earthscope-sdk.
Use backend="fdsn" when:
- the desired network is not available from EarthScope / SCEDC / NCEDC / GeoNet S3
- the archive-backed attempt fails and you want a fallback provider
- you need a non-US provider such as GEOFON, INGV, ETH, ORFEUS, etc.
This keeps the default workflow archive-first instead of HTTP-first.
The central decode path is:
- fetch raw miniSEED bytes
- decode with
pymseed - work in numpy immediately
ObsPy is not required for this path.
Once you have a TraceBundle, you can convert to:
xarray.Datasetfor labeled arrays and ML/data workflowszarrfor chunked cloud/local persistence- ObsPy
Streamfor classical seismology tooling - Earth2Studio adapters for sparse sensor and digital twin/data assimilation workflows
Existing seismic client workflows often couple:
- transport
- decode
- metadata
- downstream object model
seisfetch separates those concerns:
- transport: S3 or HTTP
- decode: miniSEED -> numpy
- output: xarray / zarr / ObsPy / Earth2Studio only when requested
That makes it useful for:
- cloud-native waveform mining
- sparse sensor ingestion
- foundation-model training pipelines
- Earth2Studio interoperability
- data assimilation and digital twin workflows
ObsPy is an optional extra here, not a dependency of the data path. The
evaluation behind that choice — including whether it changes the science — is
written up in
docs/noisepy-obspy-replacement-report.md,
with every number traceable to a committed JSON under benchmarks/results/ or
a test in tests/precision/.
| seisfetch | obspy stack | |
|---|---|---|
| Parse an 11 MB Steim2 channel-day | 21.4 ms | 37.2 ms |
Cold import |
0.08 s | 0.13 s |
| Parse peak memory | 27.7 MB | 52.0 MB |
| Installed footprint | 80.4 MB | 311.4 MB |
| arm64 Linux install | wheels | needs gcc |
The footprint line is the cloud argument: AWS Lambda caps a layer at 250 MB, so the ObsPy stack does not fit and seisfetch does. ObsPy publishes no linux/aarch64 wheels, so on Graviton it compiles from source; seisfetch and pymseed install from wheels.
Identical archive bytes were pushed through (A) obspy.read plus NoisePy's own
preprocess_raw and (B) seisfetch's parser plus the numpy/scipy ports in
seisfetch.contrib.noisepy_adapter, then through NoisePy's own compute_fft
and correlate. The pass criterion is bit-identity, not closeness:
| Harness | Result |
|---|---|
| Single-station CI.PASC, EN/EZ/NZ/ZZ at 40 sps | max abs diff 0.0 |
| Cross-station SCEDC x NCEDC x EarthScope at 20 sps | max abs diff 0.0 |
| dv/v stretching | same grid cell on all pairs |
The cross-station run matters because it exercises the Fourier-resample and sub-sample-alignment branches inside the chain. Tables, plots and the harnesses: benchmarks/RESULTS.md.
Scope note: this bit-identity result covers the preprocessing chain at
rm_resp=NO. Response removal is validated separately, below.
Response removal was the one ObsPy capability the NoisePy migration still
needed. seisfetch.contrib.response provides it in ~577 lines of numpy and
stdlib xml.etree — no evalresp C library, no ObsPy, no lxml, and no scipy in
that module. ObsPy has no pure-Python response evaluator (remove_response
calls compiled evalresp), so every stage was re-derived and then checked
against the compiled implementation:
| Check | Result |
|---|---|
evaluate_response(mode="full") vs compiled evalresp — both CI.PASC epochs, VEL/ACC/DISP, 1 mHz–19.9 Hz |
max rel diff 1.6e-10 |
remove_response_np vs Trace.remove_response — real 6.9M-sample Tohoku day, water_level=60, pre_filt |
6.6e-16 of peak |
| Same, CI.PASC demo hour in notebook 06 | 7.6e-16 of peak |
| Deconvolve that Tohoku day | 1.9 s vs ObsPy 3.6 s |
Two evaluation modes: mode="full" is evalresp-equivalent (all stages, analog
and digital poles/zeros, FIR/Coefficients with DC normalization and the
Decimation/CorrectionApplied phase advance). mode="paz" is the SACPZ
shortcut — 0.7–1.3 % error below 4 Hz but ~23 % by 16 Hz, since the FIR
anti-alias roll-off is unmodeled; don't use it above ~Nyquist/3.
Two deconvolution styles: remove_response_np ports ObsPy's water-level
method for drop-in equivalence, and translate_resp_np follows SeisIO.jl's
translation approach, taking stabilization from the target response's own
roll-off instead of a water level.
Defective metadata fails loudly — zero or missing gains, degenerate
normalization references, zero-sum FIR stages, polynomial and ResponseList
stages all raise with the stage number named, never a silent NaN or unity
gain. Not implemented, and raising rather than approximating: IIR
Coefficients stages with denominators, polynomial (blockette-62) responses,
ResponseList stages. Metadata is StationXML only; RESP and SACPZ files are
not parsed.
Full derivation, the conditional-A0 finding about evalresp's normalization rule, and the SeisIO comparison: docs/response-removal-design.md. Tutorial: notebooks/05_response_removal.ipynb.
from seisfetch import SeisfetchClient
client = SeisfetchClient(backend="s3_open")
bundle = client.get_numpy(
"CI",
"ABL",
channel="BHZ",
starttime="2024-01-15T00:00:00",
endtime="2024-01-15T00:10:00",
)
print(bundle.ids)
arrays = bundle.to_dict()NZ auto-routes to the GeoNet open-data bucket. GeoNet channels carry a
numeric location code, so pass location=:
bundle = SeisfetchClient(backend="s3_open").get_numpy(
"NZ",
"WEL",
location="10",
channel="HHZ",
starttime="2022-01-02T00:00:00",
endtime="2022-01-02T00:10:00",
)Requires earthscope-sdk and an EarthScope account that has been granted the
s3-miniseed role. See EarthScope Credentials.
client = SeisfetchClient(backend="s3_auth")
bundle = client.get_numpy(
"IU",
"ANMO",
location="00",
channel="BHZ",
starttime="2024-01-15T00:00:00",
endtime="2024-01-15T00:01:00",
)Works for any logged-in EarthScope account, including those without direct S3
access. Good for TA, IU, US, UW, _GSN, etc.:
client = SeisfetchClient(backend="fdsn", providers="EARTHSCOPE")
bundle = client.get_numpy(
"TA",
"034A",
channel="BHZ",
starttime="2010-06-01T00:00:00",
endtime="2010-06-01T00:05:00",
)
print(bundle.ids) # ['TA.034A..BHZ']
print(bundle.to_dict()[bundle.ids[0]].shape)get_stations() queries fdsnws-station and auto-routes to SCEDC, NCEDC, or
EarthScope based on the network code. No ObsPy required.
rows = client.get_stations(
"TA", channel="BHZ",
starttime="2010-06-01", endtime="2010-06-02",
)
for r in rows[:3]:
print(r["Network"], r["Station"], r["Channel"], r["Latitude"], r["Longitude"])client = SeisfetchClient(backend="fdsn", providers="GEOFON")
bundle = client.get_numpy(
"GE",
"BKB",
channel="BHZ",
starttime="2011-03-11T06:00:00",
endtime="2011-03-11T06:05:00",
)ds = client.get_xarray(
"CI",
"ABL",
channel="BHZ",
starttime="2024-01-15T00:00:00",
endtime="2024-01-15T00:10:00",
)from datetime import datetime
from seisfetch import SeismicDataFrameSource
df_source = SeismicDataFrameSource(bundle)
df = df_source(datetime(2024, 1, 15), list(ds.data_vars))from seisfetch import bundle_to_metadata_table, to_zarr, write_metadata_csv
metadata_table = bundle_to_metadata_table(bundle)
to_zarr(bundle, "quickstart.zarr", metadata=metadata_table)
write_metadata_csv(metadata_table, "quickstart.zarr")You do not need every dependency for every workflow.
Use this when you only want miniSEED -> numpy from S3:
pip install seisfetchIncludes:
numpyboto3pymseed
Use this if you want direct HTTP fallback providers:
pip install "seisfetch[fdsn]"Adds:
httpx
Use this if you need EarthScope archive access:
pip install "seisfetch[auth]"
pip install earthscope-cliAdds:
earthscope-sdk- EarthScope CLI login flow
Use this for labeled arrays and ML pipelines:
pip install "seisfetch[xarray]"Use this if you want canonical metadata tables, CSV sidecars, or metadata-aware zarr output:
pip install "seisfetch[pandas]"Use this for chunked persistent stores:
pip install "seisfetch[zarr]"Use this only if you need ObsPy interop or ObsPy-backed FDSN behavior:
pip install "seisfetch[obspy]"| Need | Install |
|---|---|
| S3 open data -> numpy only | pip install seisfetch |
| Metadata table / metadata.csv export | pip install "seisfetch[pandas]" |
| Archive-first + FDSN fallback | pip install "seisfetch[fdsn]" |
| EarthScope + SCEDC + NCEDC | pip install "seisfetch[auth]" and pip install earthscope-cli |
| xarray / ML / Earth2Studio-style workflows | pip install "seisfetch[xarray]" |
| zarr persistence | pip install "seisfetch[zarr]" |
| ObsPy interop | pip install "seisfetch[obspy]" |
| Most common research stack | pip install "seisfetch[fdsn,auth,xarray,zarr,obspy]" |
For scripts and lightweight pipelines:
pip install seisfetchFrom source:
git clone https://github.com/Denolle-Lab/seisfetch
cd seisfetch
pip install .For notebook work and development:
git clone https://github.com/Denolle-Lab/seisfetch
cd seisfetch
pixi install
pixi install -e notebooks
pixi run -e notebooks kernel-installThe notebook environment is the intended Jupyter environment for this repo.
EarthScope has two access tiers:
- FDSN web service (
backend="fdsn", providers="EARTHSCOPE") — available to any logged-in account. No special role required. - Direct S3 (
backend="s3_auth") — requires thes3-miniseedIAM role to be granted on your account. This is faster and cheaper when you are running inus-east-2.
pip install "seisfetch[auth]"
pip install earthscope-cli
es loginThe [auth] extra already installs earthscope-sdk, so there is no separate SDK install step.
# 1. Confirm you are logged in
es user get-profile # prints your name, email, institution
# 2. Confirm direct-S3 role is granted (only needed for backend="s3_auth")
es user get-aws-credentials # prints temporary AWS keys, or an errorA response of "You are not allowed to assume role 's3-miniseed'" or
UnauthorizedError means your account is logged in but direct S3 is not
enabled yet. Use backend="fdsn", providers="EARTHSCOPE" in the meantime
and email data-help@earthscope.org to request the s3-miniseed role.
from earthscope_sdk import EarthScopeClient
with EarthScopeClient() as client:
print(client.user.get_profile())
try:
creds = client.user.get_aws_credentials()
print("S3 role granted:", creds.aws_access_key_id[:8])
except Exception as exc:
print("S3 role NOT granted:", exc)es user get-refresh-token
export ES_OAUTH2__REFRESH_TOKEN="<your-refresh-token>"SeisfetchClient
|
+- get_raw() -> raw miniSEED bytes
+- get_numpy() -> TraceBundle (numpy)
+- get_xarray() -> xarray.Dataset
+- get_waveforms() -> ObsPy Stream
|
+- backend="s3_open"
| +- SCEDC open bucket
| +- NCEDC open bucket
| +- GeoNet open bucket
| +- auto-routing by network code
|
+- backend="s3_auth"
| +- EarthScope S3 via earthscope-sdk credentials
|
+- backend="fdsn"
| +- single-provider HTTP client
| +- multi-provider fan-out client
|
+- backend="obspy_fdsn"
+- ObsPy-backed fallback for harder FDSN cases
The package includes adapters in seisfetch.earth2 for Earth2Studio-style usage:
SeismicDataSource— wraps a bundle you already fetchedSeismicDataFrameSource— sparse sensor table;auto_coords=Truefills station lat/lon from the FDSN station serviceSeisfetchLiveSource— time-indexedDataSourcethat fetches on demandbundle_to_earth2
These are intended for:
- sparse sensor tables
- observation pipelines
- foundation-model data preparation
- Earth2Studio / digital twin workflows
Typical path:
miniSEED -> numpy -> xarray / sparse dataframe -> Earth2Studio adapter
SeisfetchLiveSource has the shape every other Earth2Studio source (GFS,
ERA5, ...) has: you call it with timestamps and it fetches, auto-routing per
network across all four archives and caching day bundles in memory.
from datetime import datetime
from seisfetch.earth2 import SeisfetchLiveSource
source = SeisfetchLiveSource(
channels=["CI.PASC..BHZ", "BK.PKD.00.BHZ", "II.PFO.00.BHZ"],
window_s=3600,
calibrate="gain",
)
da = source(datetime(2022, 1, 2, 6)) # -> (time, variable, sample) DataArrayChannels are NET.STA.LOC.CHA strings; the returned variable coordinate
spells them with underscores (CI_PASC__BHZ), which is also what the
optional variable= argument accepts. All channels in one call must share a
sampling rate — request mixed rates (a 40 sps BHZ alongside a 100 sps
HHZ) in separate calls.
Physical units are required — this source never returns raw counts:
calibrate="gain"(default): divide by the channel's total sensitivity from the FDSN station service. Exact at the reference frequency, one metadata request per channel, no extra dependencies.calibrate="response": full spectral deconvolution throughseisfetch.contrib.response(StationXML fetch plus an evalresp-equivalent evaluator). Agrees with ObsPy to 7.6e-16 of peak amplitude on real data.
fetch() is genuinely async (asyncio.to_thread), so pipelines can overlap
this source with others.
A few common end-to-end tasks the package is designed for.
from seisfetch import SeisfetchClient
bundle = SeisfetchClient(backend="s3_open").get_numpy(
"CI", "ABL", channel="BHZ",
starttime="2024-01-15T00:00:00",
endtime="2024-01-16T00:00:00",
)
data = bundle.to_dict()["CI.ABL..BHZ"] # int32 numpy array, full dayrequests = [
{"network": "CI", "station": s, "channel": c,
"starttime": "2025-04-14T17:07:30", "endtime": "2025-04-14T17:10:30"}
for s in ["ABL", "SDD", "PASC"]
for c in ["BHZ", "BHN", "BHE"]
]
client = SeisfetchClient(backend="s3_open")
summary = client.get_numpy_bulk(requests, max_workers=8)
print(summary.succeeded, "/", summary.total)Missing station-channel combinations are reported in summary.failed and do
not abort the run.
from seisfetch import SeisfetchClient, route_network
def get_archive_first(net, sta, *, starttime, endtime, location="*", channel="*",
fallback="GEOFON"):
dc = route_network(net)
primary = "s3_open" if dc in {"scedc", "ncedc"} else "s3_auth"
try:
return SeisfetchClient(backend=primary).get_numpy(
net, sta, location=location, channel=channel,
starttime=starttime, endtime=endtime,
)
except Exception:
return SeisfetchClient(backend="fdsn", providers=fallback).get_numpy(
net, sta, location=location, channel=channel,
starttime=starttime, endtime=endtime,
)from seisfetch import bundle_to_metadata_table, to_zarr, write_metadata_csv
metadata = bundle_to_metadata_table(bundle)
to_zarr(bundle, "package.zarr", metadata=metadata)
write_metadata_csv(metadata, "package.zarr")The resulting package.zarr/ contains channel groups plus a
metadata/channel_table group readable with xarray.open_zarr(..., group="metadata/channel_table").
from datetime import datetime
from seisfetch import SeismicDataFrameSource, bundle_to_xarray
ds = bundle_to_xarray(bundle)
df = SeismicDataFrameSource(bundle)(datetime(2024, 1, 15), list(ds.data_vars))
df[["time", "variable", "network", "station", "channel",
"sampling_rate", "amplitude_rms", "num_samples"]].head()Examples:
# SCEDC open S3 (no auth)
seisfetch download CI ABL -s 2024-01-15 -e 2024-01-15T01:00:00 -c BHZ -o data.mseed
seisfetch numpy CI SDD -s 2024-06-01 -c BHZ -o data.npz
seisfetch zarr CI ABL -s 2024-01-15 -c BHZ -o data.zarr
# EarthScope (requires `es login` and the s3-miniseed role)
seisfetch download IU ANMO -s 2024-01-15 -e 2024-01-15T01:00:00 -c BHZ --backend s3_auth -o anmo.mseed
# Routing and provider info
seisfetch info --route CI
seisfetch info --providers
# Bulk
seisfetch bulk requests.csv -o output/ -f npzSee notebooks/ for worked examples:
- 01_quickstart.ipynb
- 02_bulk_mining.ipynb
- 03_xarray_zarr_pipeline.ipynb
- 04_earth2studio_interop.ipynb
- 05_response_removal.ipynb — instrument response removal without ObsPy
- 06_cross_correlation_three_archives.ipynb — four stations, three archives, two months of NoisePy cross-correlations
Notebook setup instructions are in notebooks/README.md.
pixi run test
pixi run test-cov
pixi run test-int| Package | Status | Role |
|---|---|---|
numpy |
core | array container |
boto3 |
core | S3 transport |
pymseed |
core | miniSEED decode |
httpx |
optional [fdsn] |
FDSN HTTP client |
pandas |
optional [pandas] |
canonical metadata tables and CSV export |
xarray |
optional [xarray] |
labeled dataset output |
zarr |
optional [zarr] |
persistent chunked storage |
obspy |
optional [obspy] |
ObsPy interop and alternative FDSN backend |
earthscope-sdk |
optional [auth] |
EarthScope S3 credentials |
See THIRD_PARTY_NOTICES.md for attribution and licenses.
Benchmarks with plots: benchmarks/RESULTS.md (or the self-contained RESULTS.html). Changes: CHANGELOG.md. Response-removal tutorial: notebooks/05_response_removal.ipynb.
When using data accessed through seisfetch:
- EarthScope: cite the network operators and NSF SAGE facility
- SCEDC: doi:10.7909/C3WD3xH1
- NCEDC: doi:10.7932/NCEDC
- GeoNet: GNS Science, GeoNet open data (CC BY 4.0) — cite per GeoNet's data policy
- other FDSN providers: cite the underlying network/provider
Software references:
pymseed/libmseed: EarthScope Data Services- NoisePy S3 access pattern: Jiang & Denolle (2020), doi:10.1785/0220190364
- ObsPy: Beyreuther et al. (2010), doi:10.1785/gssrl.81.3.530
MIT AND LGPL-3.0-only. The package is MIT (LICENSE) with one
exception: seisfetch/contrib/obspy_ports.py
contains numerically exact translations of ObsPy routines and is
LGPL-3.0-only (ObsPy is © The ObsPy Development Team, LGPL v3). Using
seisfetch as a library is unaffected; redistributors of modified versions of
that one file must honor LGPL terms. Details in
THIRD_PARTY_NOTICES.md.