← All work

Auction price prediction

Prices an auction lot from its recorded attributes, and returns a range showing how confident the model is.

Role
ML engineer
Year
2026
Status
internal
Built with
Python, TensorFlow, Keras, DistilBERT, AWS SageMaker
Context
GoPrime Systems, production auction platform
  • 850 → 92MSE, earlier build
  • 68% (1σ)Prediction interval
  • 9Categorical features
  • 6Numeric features

The model and the sale history it trains on belong to my employer, so the source stays private. What is described here is the architecture and the decisions behind it.

An auction platform takes in items faster than anyone can value them by hand. The model prices each lot from what is already known about it: category, condition, location, weight, make, material, how many items are in the lot, and a handful of engineered features on top.

The first version was a hybrid: DistilBERT over the item text, a convolutional path over the photographs, and the structured fields alongside both. It brought mean squared error down from 850 to 92, which was low enough to make automatic lot bundling worth doing.

The model is different now. What follows is what changed and why.

From a number to a range

The model predicts a distribution over a price rather than a price.

The network ends in two heads instead of one: a mean and a log variance. They are concatenated into a single tensor so the served signature still has one output, and training uses Gaussian negative log likelihood instead of mean squared error.

mu = y_pred[:, 0]
log_var = tf.clip_by_value(y_pred[:, 1], -4.0, 10.0)
return 0.5 * tf.reduce_mean(log_var + tf.square(y_true - mu) / tf.exp(log_var))

Under squared error the model is rewarded for being right on average and has no way to report that it does not know. A common item in a familiar category and a one-off with a make nobody has seen get the same confident treatment.

Under this loss there are two ways to reduce the error: predict better, or report more uncertainty. Widening the interval reduces the squared error term but costs directly in the log_var term, so claiming uncertainty everywhere costs more than it saves. The minimum sits where the predicted spread matches the error the model actually makes, which is what makes the interval usable.

The clip on the log variance keeps training stable. Without it, a handful of items the model finds baffling drive the variance towards infinity, where the loss is finite and the gradient is not.

MSE is no longer the headline number. Negative log likelihood is not readable as a quality measure, so mean absolute error on the mean head is tracked alongside it for monitoring. The loss trains the model, the MAE says how it is doing.

Getting the interval into the right space

The target is log-transformed before training.

Prices are multiplicative. The gap between a $20 item and a $40 item is the same kind of gap as between $2,000 and $4,000, and a model trained on raw dollars spends its capacity on the expensive tail while treating everything cheap as approximately zero. Training on log price fixes that, and it means the variance the model learns is a variance in log space.

So the interval is exponentiated back rather than added in dollars:

sigma = np.sqrt(np.exp(log_var))
prices  = np.expm1(mu)
lower_68 = np.expm1(mu - sigma)
upper_68 = np.expm1(mu + sigma)

The result is asymmetric. A prediction of $1,250 comes back with a range of roughly $820 to $1,900, less room below than above. A symmetric dollar interval would put the lower bound below zero on cheap items and understate how far an expensive one can run.

The spread is also reduced to a single confidence score by taking exp(-sigma), which lands between zero and one. A sigma of 0.1 gives about 0.90, a sigma of 0.7 gives 0.50. An operator does not want to reason about log-space standard deviations, and a number that behaves like a percentage is something a user interface can act on.

What got switched off

The image path is gone and the text path is dormant.

DistilBERT is still in the model code behind a configuration flag, with mean pooling over non-padding tokens rather than the CLS token because the text fields here are short. The two-phase training schedule that goes with it also survives: train the dense head with the encoder frozen, then unfreeze and fine-tune end to end at a much lower rate. Turning it back on is one flag.

Keeping a disabled path around has a cost, since dormant code rots. I kept it because the schedule took experimentation to get right and because the payload contract already carries the text fields.

Backwards compatible on purpose

The model changed shape, gained an output and changed its loss, and the application consuming it did not have to change anything.

The endpoint accepts three payload shapes: a wrapped batch array, a single bare object which is auto-wrapped, and a plain array. Categorical identifiers are accepted as integers, as strings, or as float-formatted strings, and normalised internally, so 1, "1" and "1.0" all mean the same thing. Unknown category values fall back to a default index instead of raising. Unrecognised fields in the payload are ignored rather than rejected. Text fields are still accepted even though nothing currently reads them.

The serving handler reads both model shapes. If the artifact returns two columns it produces the mean, the interval and the confidence score; if it returns one, it produces the price alone and reports zero confidence. Rolling back to the previous artifact does not require rolling back the handler.

Response shapes are preserved too, down to the nesting the C# client expects.

The reason for all of it is that a model improvement should not require a coordinated release across two teams. The endpoint gained a capability, every existing caller kept working, and the application can adopt the interval when it is ready.

Saying when not to trust it

Alongside the prediction, each instance comes back with a list of validation warnings and an unreliable flag: an unfamiliar category value, a numeric feature well outside the range seen in training, text that would have been truncated.

The uncertainty head covers what the model knows it does not know. The warnings cover the case where the input itself is outside what the model was trained on, which the variance does not always catch.

Where it is

Deployed on SageMaker behind a TensorFlow inference container, serving the platform’s intake.

The consuming application still reads the point estimate and discards the range, because it has not been updated yet. So the most useful thing the model now does, flagging when a lot should go to a person instead of being priced automatically, is sitting unused in the response. The compatibility work above is what keeps that a one-sided change: the capability is live, and adopting it only requires work on the application side.