10.1. Stratagem Concepts#

A Geometrics/EMI Stratagem receiver does not write EDI files. It writes its own per-station, per-component ASCII tables of raw cross-spectral values, and getting from there to something pycsamt.emtools or an inversion engine can read requires a conversion step the instrument vendor never automated: WinGLink, a Windows desktop program run by hand, once, outside pyCSAMT. Everything in pycsamt.stratagem exists on one side or the other of that conversion – reading the raw files for hardware-level diagnostics before WinGLink ever runs, or correcting and finishing the EDI files WinGLink produces. This page covers the raw format and the WinGLink boundary; the pages that follow pick up the pipeline from each side of it.

10.1.1. Raw Hardware File Format#

A Stratagem delivery names its component files by prefix and extension rather than by a descriptive stem: X2HX.001, X2HX.002, … for the Ex electric-field component, Y2HX.001, … for Ey, Z2HX.001, … for Hz. Every file for a given component shares the same stem (X2HX) – the station number lives in the extension, not the name. Opened as text, each row is a whitespace-separated record:

Column

Field

Meaning

0

Frequency (Hz)

One hardware frequency bin.

1

Instrument constant

Fixed per Stratagem unit, ~2.93 in every K2 file.

2

Stack count

Time-series windows averaged into this bin; 0 means no usable measurement.

3-18

Cross-spectral values

16 real numbers – real/imaginary pairs across the quantities the instrument correlates, before WinGLink turns them into the impedance tensor components an EDI’s >ZXXR/>ZXXI blocks store.

StratagemRawReader parses exactly this table. It does not compute an impedance tensor – that is still WinGLink’s job – it only turns columns 0 and 2 into a frequency grid and a boolean mask, which is already enough to see how much of a station’s recording is usable before waiting on a full WinGLink run:

>>> from pycsamt.stratagem import StratagemRawReader

>>> rdr = StratagemRawReader("data/stratagem/K2/k2-HX", component="X").fit()
>>> rdr.n_stations_, rdr.n_freqs_
(87, 292)
>>> rdr.freqs_[:5]
array([ 6.25,  7.5 ,  8.75, 10.  , 11.3 ])
>>> rdr.freqs_[-3:]
array([ 97500.,  98800., 100000.])
>>> round(rdr.station_coverage(), 3)
0.794

292 frequency bins spanning 6.25 Hz to 100 kHz is the fixed table this particular Stratagem unit records at every station; station_coverage() folds the whole (87 station, 292 frequency) mask down to one number, the fraction of cells where the stack count was non-zero. That average hides which stations are actually the problem, which station_frame() answers directly:

>>> cols = ["station", "usable_freqs", "coverage", "med_stacks"]
>>> rdr.station_frame().sort_values("coverage")[cols].head()
     station  usable_freqs  coverage  med_stacks
59  X2HX.060           106  0.363014        36.0
58  X2HX.059           106  0.363014        33.0
57  X2HX.058           106  0.363014        30.0
56  X2HX.057           106  0.363014        33.0
55  X2HX.056           106  0.363014        36.0

Stations 56 through 61 all land on exactly the same low coverage, which is itself informative – a handful of scattered bad frequencies would scatter across many stations, but a flat run of six neighbouring stations sharing one coverage value points at something that affected the whole occupation there (cultural noise, poor ground coupling, whatever the field crew logged for that stretch of line), not at isolated hardware glitches. Plotting the full mask makes the same pattern visible across the whole line at once:

>>> fig = rdr.plot_coverage(
...     kind="snr", title="K2 line -- hardware SNR coverage (X component)",
... )
>>> fig.savefig("concepts_raw_coverage.png", dpi=170, bbox_inches="tight")
Station-by-frequency SNR coverage heatmap for the K2 line, X component.

Green marks a usable (station, frequency) cell, red an absent one. The solid red band across stations 56-61 is exactly the low-coverage run surfaced above. Below roughly 80 Hz the mask alternates column by column rather than failing outright – every station loses the same scattered low-frequency bins, consistent with a shared limitation of the receiver’s low-frequency response rather than a per-station fault. A thin red column at the very top of the frequency range shows the same thing happening at the opposite, high-frequency edge of the band.#

10.1.2. Station Numbering and Natural Sort#

Because the station number sits in the file extension, sorting these files the way a filesystem or a plain sorted() call would is actively wrong once station numbers stop sharing the same digit width. Both X2HX.001 and X2HX.087 compare as strings starting with the identical stem, so the comparison falls through to the extension – and string comparison of "1" against "10" puts "10" first, because it compares character by character rather than by numeric value:

>>> from pathlib import Path
>>> from tempfile import TemporaryDirectory
>>> from pycsamt.stratagem.io import _station_number

>>> with TemporaryDirectory() as tmp:
...     d = Path(tmp)
...     for sid in [1, 2, 10, 87]:
...         (d / f"X2HX.{sid}").touch()
...     plain = sorted(d.glob("X2HX.*"))
...     by_number = sorted(d.glob("X2HX.*"), key=lambda p: _station_number(p.name))
...     print([p.name for p in plain])
...     print([p.name for p in by_number])
['X2HX.1', 'X2HX.10', 'X2HX.2', 'X2HX.87']
['X2HX.1', 'X2HX.2', 'X2HX.10', 'X2HX.87']

StratagemRawReader always sorts by the value its private _station_number helper extracts from the extension, never by the raw path string, which is what natural sort means in this context. The K2 raw files happen to be zero-padded (001087), so plain and natural order agree for this particular delivery – the demonstration above uses un-padded numbers on purpose, because that is exactly the case a delivery without zero-padding would hit, and there is no guarantee every Stratagem export is padded consistently. EDIBatch applies the same natural-sort key, via its own _edi_sort_key helper, when it loads a directory of WinGLink EDI exports, for the same reason.

10.1.3. Component Format Inconsistencies#

The 19-column table above describes the X and Z component files. The K2 delivery’s Y files share the same directory and the same naming convention, but not the same content:

>>> from pathlib import Path

>>> x_bytes = Path("data/stratagem/K2/k2-HX/X2HX.001").read_bytes()
>>> y_bytes = Path("data/stratagem/K2/k2-HX/Y2HX.001").read_bytes()
>>> len(x_bytes), len(y_bytes)
(61904, 1573088)
>>> x_bytes[:60]
b' 6.250e+000 2.930e+000 0.000e+000 0.000e+000 0.000e+000 0.00'
>>> y_bytes[:60]
b"\x04\x00\x00\x00@\x05\xd0\x07\xd0\x07\x000\x04\x00\x07\x01y\t\x1d\x00\xc6\xf6S\x01\xb6\x06\x18\x01y\xf8\xbe\xfe\xa8\xfe\xc7\xfea\x00j\xfc)\xf9`\xfcg\x068\xfd;\xfb\xc6\xfcX\x05\xcd\x01\xbf\x04'\x01"

Y2HX.001 is 25 times larger than its X counterpart and opens as binary, not the documented ASCII table – a raw time-series capture rather than the 19-column spectral summary. Nothing in the file name distinguishes the two, so it is worth checking before trusting component='Y' against a new delivery, because StratagemRawReader does not detect the mismatch and fails silently rather than raising:

>>> rdr_y = StratagemRawReader("data/stratagem/K2/k2-HX", component="Y").fit()
>>> rdr_y.n_stations_, rdr_y.n_freqs_
(87, 4093)
>>> rdr_y.freqs_[:5]
array([6., 8., 4., 7., 0.])

292 real frequency bins became 4093 – the regex numeric extractor happily pulls digit-looking sequences out of whatever text the binary bytes decode to, and a frequency grid of [6., 8., 4., 7., 0.] is the result. Nothing raises, because nothing about the file is malformed ASCII from the parser’s point of view; it is simply the wrong kind of file. The safe reading is component-by-component: confirm a handful of files in a new delivery actually look like the table above before trusting the mask it produces, rather than assuming every X/Y/Z triplet was captured the same way.