← All work

KYC document verification

Checks that the face on an ID belongs to the person submitting it, and that their proof of residence says where they claim to live.

Role
ML engineer
Year
2024
Status
archived
Built with
Python, DeepFace, EasyOCR, fuzzywuzzy
Context
GoPrime Systems, internal project

Built for an employer, so the source is not mine to publish. I wrote the ML pipeline; a colleague later moved it into a Flask and Redis service, which is why the repository history sits with him. The product it was built for never went live.

Onboarding a customer means answering two questions about two documents.

The first is whether the person submitting the application is the person on the identity document. The second is whether the utility bill or bank statement they uploaded shows the address they typed into the form. One is a face problem and one is a text problem, and most of the difficulty is in the second.

Matching an address against a blob of OCR

Reading a proof of residence gives you one long run of text with no structure. The document is a bank statement or a council bill, so the address is in there somewhere, surrounded by account numbers, dates, tariffs and marketing copy. You know what you are looking for but not where it is.

Two obvious approaches fail. Substring matching breaks on the first OCR error, and there is always an OCR error: a scanned street name comes back with a zero for an O or a missing space. Fuzzy matching the address against the whole extracted text fails the other way, because similarity against a thousand-character blob is dominated by the characters you do not care about. A perfect match scores badly because the haystack is large.

So the matcher slides a window. It takes the target string, measures it in words, and walks a window of that many words across the OCR output, scoring each position and keeping the best. A three-word street name is only ever compared against three-word spans, so the score means what it looks like it means.

Around that is a cheap-first ordering. The address on file is split into lines and each is checked independently, first by plain substring containment, and only if that fails by the sliding fuzzy match against a similarity threshold of 90. Most lines on most documents match exactly and cost nothing.

The output is not a boolean. It returns which elements of the address failed to verify, so a reviewer or a support agent knows whether the applicant mistyped a house number or uploaded a bill for a different property. A rejection that says only “denied” turns into a support ticket.

Photographs taken by people holding phones

The face check compares the portrait on the ID against the submitted selfie using a pretrained face recognition model and cosine distance. That part is mostly picking sensible defaults.

The one addition is that the pipeline retries the comparison at four orientations. People photograph their ID card on a table, sideways, in whatever rotation the phone recorded, and face detection on a sideways face tends to fail confidently rather than gracefully. Rotating through 0, 90, 180 and 270 degrees and accepting the first orientation that verifies costs nothing when the image is upright, because that is the first attempt, and rescues the submission when it is not.

From pipeline to service

I built the above as a self-contained pipeline. A colleague later moved it into a Flask endpoint that enqueues to Redis with a worker draining the queue, which suits the workload: the models load real weights, so a synchronous request would spend its life in cold starts and timeouts, and the imports belong in the worker rather than in a web process whose only job is to accept a job.

One behaviour did not survive the port. The rotation retry did not make it across, so the service version calls the face comparison once, at whatever orientation the image arrived in. Nothing about the refactor was wrong, and the dropped behaviour looks like an implementation detail until a user submits a sideways photograph.

The bug I would fix first

The verdict logic has a defect, and the flow is mine.

Each check writes an overall status. The face check writes a denial if it fails. The address check then writes its own verdict afterwards, unconditionally. So a submission where the face does not match but the address does is reported as verified, because the second write overwrites the first. The individual field recording the face result stays false, so a careful consumer could still catch it, but the field that reads as the verdict is wrong.

Two checks that must both pass should be combined once at the end rather than each writing the shared answer as it finishes. It never mattered in practice because this never carried real applicants.

Where it went

Nowhere. The product it was built for never launched, so the pipeline was finished, deployed to a serverless endpoint for testing, wrapped into a service, and then shelved along with the rest of it.

The address matcher is the piece I would reuse. Matching a known short string against unstructured extracted text is not specific to KYC, and every OCR pipeline runs into it eventually.