← All work

multiresolution-transformer

Forecasts market movement by reading several timeframes of the same series inside one attention stack.

Role
Sole engineer
Year
2026
Status
in development
Built with
Python, TensorFlow, Keras, Polars
Go look
  • 200GB+Raw tick data
  • 10TB+Memory explosion avoided

Given the last 64 five-minute candles and the last 64 hourly candles, the model predicts two numbers: the highest and the lowest price over the next twelve five-minute candles. Together they bound the next hour’s range.

The finance is the application rather than the point. What follows is the machinery, because on a problem this noisy the machinery decides whether a result means anything. The result up front: the model does not beat a random walk out of sample. The rest is what the pipeline does and what it cost to find that out.

Gated blocks with stochastic depth, written from scratch

The stack is four transformer blocks. StochasticGatedTransformerBlock adds three things to the usual attention-plus-feedforward pattern.

A learned gate on each sublayer output. A sigmoid Dense produces a per-token, per-channel gate that multiplies the attention output, and another does the same for the feedforward output. Instead of every token receiving its sublayer’s full contribution, the block decides per token how much to let through. With two resolutions sharing one sequence, an hourly token and a five-minute token are asking different questions of the same attention, and a fixed residual gives them the same answer.

The gate reads a different tensor than attention does. Attention runs over a noised copy of the input, and the gate is computed from the clean input:

attn_output = self.att(query=inputs_with_noise, key=inputs_with_noise, value=inputs_with_noise)
gate_val = self.gate_att(inputs)

The noise stops the block memorising exact configurations. Letting it reach the gate as well would mean the decision about how much to trust a sublayer is itself made on corrupted evidence.

Stochastic depth. With probability stochastic_depth_rate the entire block is skipped during training and the input passes through untouched, via a tf.cond on the training flag. The rate ramps from 0.05 in the first block to 0.2 in the last, so early layers are nearly always present and later ones learn to be optional. Four blocks that must all fire is a deeper commitment than the data supports.

Noise in three places, all training-only. Additive Gaussian on the block input at stddev 0.01, multiplicative noise on both gate values at mean 1.0 and stddev 0.1, and a feature-dropout mask on the feedforward hidden activations rescaled by 1.1 to preserve magnitude. Each is wrapped in a tf.cond on training, so inference is deterministic and the saved model behaves the same on every call.

That is a lot of simultaneous regularisation, and I never ran the ablation that would say which parts earned their place. Training was stable across long runs with no divergence, which is what I was after, but stable is not the same as justified piece by piece.

Normalisation as a pipeline rather than a step

Price levels are not comparable. GBPUSD trades near 1.27 and gold near 2000, so a single model can only see both if prices stop being prices first.

Every value is rescaled against the high and low of the preceding 144 candles:

x_normalized = (x - window_min) / (window_max - window_min)

The formula is simple. What has to be true around it is not.

The rolling extremes are computed in column blocks. Materialising 144 shifted copies of the frame at once is the obvious implementation and it is enormous. normalize_by_window instead walks the lookback in blocks of twenty, shifting that many columns, folding them into a running window_max and window_min, then dropping them before the next block. Peak memory is set by the block size rather than the window length.

Labels use the previous row’s window. A label describes the future. If it were normalised against a window that includes that future, the target would carry information about itself. So _normalized_for_label columns are scaled by the shifted bounds. It is a one-line change that produces excellent and meaningless results if you get it wrong.

The bounds are carried through to the other end. norm_window_min, norm_window_max and the raw close are written into the chunk CSVs, so any prediction converts back to a real price at evaluation time. Normalised error is not meaningful on its own, because its scale depends on how volatile the window happened to be, so the same MAE means different things in different regimes. Everything is reported both ways.

The incomplete hour gets its own treatment. At 14:25 the 14:00 hourly candle does not exist yet. Waiting for the close throws away 55 minutes; filling it in leaks. Instead the running open, high, low and close so far are normalised on the hourly scale, through a backward merge_asof onto the hour bounds, so the partial candle is comparable to the closed ones beside it. Two scalars tell the model what it is looking at: minutes into the hour, cyclically encoded, and the partial hour’s length. A five-minute-old hour and a fifty-five-minute-old one are different objects.

Splitting overlapping windows without leaking

Consecutive samples share 63 of their 64 candles. A shuffled split puts near-copies of training rows into the test set and returns meaningless scores.

split_multiresolution_chunks cuts the series into large contiguous chunks and splits train, validation and test within each chunk in time order, never shuffling. The heads of validation and test are then trimmed by the lookback length, so their earliest samples cannot reach back into training rows.

That much is standard. The part specific to two resolutions:

hour_start_pos = np.searchsorted(hour_times, chunk_start_time, side='right') - 1
earliest_hour_idx = max(0, hour_start_pos - hour_lookback + 1)
earliest_hour_time = hour_times[earliest_hour_idx]
safe_start_idx = np.searchsorted(min5_times[start_idx:end_idx], earliest_hour_time, side='left') + start_idx
start_idx = max(start_idx, safe_start_idx)

Find the hourly index at the chunk boundary, walk back the full hourly lookback, then come back into the five-minute index to find the first row whose hourly history stays inside the chunk. Trim the chunk to start there.

Without it the five-minute split looks clean while the hourly branch reaches into the previous chunk, which is now the test set, and every check you would naturally run still passes. That is why the function is named for the multiresolution case rather than being a generic splitter.

The window is where transformation belongs

Upstream of everything above sits north of 200GB of raw ticks across nineteen instruments, which preprocessing collapses into the derived OHLC chunks the pipeline reads, tens of gigabytes on disk. Even that does not fit in memory, and since consecutive samples overlap by 63 of 64 candles, a materialised window tensor would have multiplied it by the lookback again and run into the terabytes. So nothing is precomputed. InstrumentChunkManager holds a bounded number of chunks under an LRU cache and evicts the rest, and windows are cut on demand.

extract_sample builds one training example: it slices the primary window out of that cached chunk, then locates each secondary resolution by binary search on the sample’s timestamp.

Any window-local transformation slots into the same place. The augmentation already lives there:

if self.config.add_noise_5min:
    main_input = add_gaussian_noise(main_input, self.config.noise_std_5min, self.config.noise_probability_5min)

add_gaussian_noise applies its noise on a linear gradient across the sequence: the oldest candle takes the full standard deviation, the newest takes none. Distant history is a rough signal about regime; the recent candles are what the label depends on and they stay sharp.

Each secondary resolution carries its own noise settings, so an hourly window can be augmented differently from a five-minute one. SecondaryResolution entries are a list, matched by binary search with a per-instrument time threshold that guarantees every resolution has a full lookback before a sample is emitted.

The transformation that did not survive

That hook exists in that shape because I wanted to put Multivariate Empirical Mode Decomposition into it. Each window would arrive already decomposed into intrinsic mode functions, handing the model a separated view of the signal instead of asking attention to find the structure itself.

It never shipped, because of cost. MEMD is iterative sifting with envelope interpolation across many projection directions, and the recommended direction count scales with the number of channels. Run per sample, per epoch, against windows cut on demand, it stopped being a preprocessing step and became the training loop.

I got it roughly twenty times faster than where it started, and it was still the dominant cost of a step by a wide margin, so I cut it. Twenty times was a real speedup and nowhere near enough. Finding the ceiling early was cheaper than finding it after a month of training runs.

What remains is the shape. The generator still transforms per window rather than per dataset, so the next candidate can be tried in an afternoon. legacy/memd.py is kept as the third-party translation it started from, with its attribution header intact, and the legacy README records that it was explored and never wired in.

What actually came out of it

The augmentation grid search, on silver, validation loss:

noise_std noise_prob best val_loss val MAE
0.0 0.0 9.826 4.211
0.001 0.2 9.712 3.940
0.001 0.7 9.838 3.975
0.001 0.9 9.862 3.592

Augmentation helps by about one percent, and the search was cut short of the full grid.

Against the naive predictor that says the next high equals the current close, the model establishes no durable edge. Both normalised and denormalised metrics are printed side by side, because improvements that appear in normalised space frequently vanish once converted back to real prices, and the denormalised column is the one that counts.

What worked: training is stable across long runs, the multi-resolution merge trains without either branch collapsing, and the pipeline handles nineteen instruments of tick-derived data without leakage. What did not: turning any of that into out-of-sample predictive power worth acting on.

The clearest unfinished piece is that the generator emits N resolutions while the model still builds exactly one secondary branch, so a four-hour or daily stream currently has no consumer. Looping the CNN, positional encoding and type embedding over the resolution list is the next step, and where I would start if I picked this up again.