The trading problem: why market data is a deep learning challenge

ML and modelling

Large language models and alternative data often dominate conversations about deep learning in trading. At high-frequency horizons, Karun Rao looks instead at the order book, where models learn from noisy, event-based market data whose structure is hard to capture by hand.

Karun RaoQuantitative Researcher

The trading problem: why market data is a deep learning challenge

When most people hear “deep learning in finance”, they might think of large language models parsing news headlines or sentiment analysis of social media feeds. At Optiver, the main driver of our high-frequency trading strategies is market data: the raw stream of order inserts, cancellations, modifications and trades flowing from exchanges.

In high-frequency trading, prediction horizons are typically seconds to minutes. In this article, I’ll try to give some intuition for what market data looks like and explain why deep learning can be useful for finding predictive patterns within it.

The trading problem

Given everything we can observe about the market right now, what will the price do next?

image1 (1).png

The research pipeline

Historically, the trading industry spent a lot of time on feature engineering to convert raw market data into predictive signals a model can use. These signals generally describe the state of the current order book or summarize how it has evolved over time. Feature engineering can also transform non-stationary data into signals that are more stationary and easier for the model to learn from.

A model combines those features into a price forecast. If you wanted to try a new idea, you had to write the feature, regenerate the dataset, train new models from scratch and then evaluate via a backtest whether the idea adds value.

Hand-crafted features hit a ceiling

There are specific areas where we have strong conviction that information exists in the data, but hand-designing features to capture it has been very hard. Finding patterns in time-series data that go beyond simple moving averages and capture temporal context is one example. In these areas, deep learning is, to us, a better method.

I use the analogy of convolutional nets. Back when I studied AI around 2010, people used hand-designed convolutional filters as features in visual tasks. CNNs changed that approach by learning the filters as part of the model. We have parallels in market data.

Instead of treating features and models as separate steps, we can train a model that takes raw market data and produces a price forecast, with feature extraction as part of the learning process. The model learns what information it can throw away, what it should retain and how to combine it for the task at hand.

From Level 1 to Level 3 market data

To motivate why market data is an interesting modality for deep learning, it helps to look at the different levels of information available.

Level 1: order-book dynamics

With Level 1 market data, you’re given the top bid and ask price and size. A simple weighted-average price is probably one of the most useful signals at this level. Deep learning on a single snapshot like this? Probably not worth it.

image1 (2).png

Consider what happens when you view this as a time series. Three order-book snapshots that look nearly identical in isolation suddenly tell a story when you see them in sequence. The bid stays at 200 lots while the size on the offer falls from 100 to 90 and then to 80. If you had to guess the next event, you might expect it to fall again.

The dynamics can get surprisingly complex. What if the next event takes the offer from 80 lots to 82? A simple exponential moving average might still show a downward trend, but perhaps you want to weigh that counterfactual evidence more heavily.

There are a lot of ideas you can explore through feature engineering here, but every one requires a full research pipeline. The deep-learning view makes feature extraction part of the learning process.

Level 2: market depth

image1 (4).png

Level 2 adds several bid and ask prices and the aggregate size available at each one. A multi-level weighted-average price might capture a large portion of the information in a single snapshot, so as with Level 1, the more compelling problem is how the order book changes over time.

Level 3: individual orders and queues

image1 (4).png

Now scale this up to Level 3 data, where you see the price, size and timestamp of every individual order insert, cancellation, modification and trade. These events can be used to construct the full order book, including the queue of orders at each price level.

The top bid and ask can have the same aggregate size while being made up very differently. One side might be dominated by a single large order, while the other contains a broader range of smaller orders. At this point, feature engineering starts to become more difficult.

Queue position carries information too. Most markets are based on time priority, so orders inserted earlier at a given price execute before those inserted later. A 200-lot order at the front of the queue therefore contains different information from the same order at the back.

You might also notice a string of orders all of size seven, placed one after another. As a human, you may guess they came from the same market participant, although the data alone cannot establish that. But will the model make the same inference? Once the sizes differ slightly or other orders appear between them, this becomes a surprisingly hard feature to engineer by hand.

The dimensionality grows again once we move beyond a single asset. Depending on what we trade, we may have tens to thousands of correlated assets that can be used to price one another, while the full order book continues to evolve over time. Together, those relationships and the order sequence create a rich spatio-temporal dataset and a much stronger case for learned feature extraction than any single snapshot.

Why market data is difficult to model

image1 (5).png

Richer inputs create more scope for learned feature extraction, but they also make the modelling problem harder. Many of the architectures above were developed for language or vision, so applying them to market data means adapting them to four characteristics in particular.

  • Low signal-to-noise ratio: Financial data contains a large amount of noise, and a great model may have an accuracy of 52%. That might sound low in other domains, but in trading it can still be valuable. The narrow margin between signal and noise also makes overfitting a serious concern.
  • Non-stationary data: Markets are forever changing, with distributions shifting as regimes change and relationships between assets changing in fundamental ways over time. A model that works well in one period may behave differently in the next.
  • Irregular sampling: With a few exceptions, high-frequency trading data is event-based. Events may be nanoseconds apart during a busy period or seconds apart when activity is lower, so the model has to account for both their order and the changing gaps between them.
  • Long sequence lengths: A typical day can contain hundreds of millions of messages. Even with techniques such as sliding-window attention, a context window of 100,000 events might cover only a minute of wall-clock time on average and much less when the market is busy. We often need more context, making this difficult from both an engineering and research perspective.

Even a model that handles those context lengths still has to meet a production latency budget, which may require custom inference hardware designed with FPGA engineers. In high-frequency trading, every nanosecond matters.

From feature design to architecture design

It’s not that we want to use deep learning to automate the research process; we still very much make use of domain expertise. The shift is from engineering each feature by hand towards designing architectures best capable of extracting structure from the data directly.

The order book contains patterns that are genuinely hard to express in human-designed terms, including queue dynamics, cancellation behavior and short-lived relationships between correlated instruments. Whether deep learning can reliably find those patterns, and keep finding them as markets adapt, is the research question we're working on.

About the author

Karun Rao is a Quantitative Researcher at Optiver and leads deep-learning research in the US high-frequency trading team. He joined Optiver in Amsterdam in 2012 and is now based in New York. Karun holds an MSc in Artificial Intelligence from the University of Amsterdam and a BSc in Computer Science from Cornell University.

Related articles

View all

A conversation with NVIDIA on how accelerated computing is changing research and engineering in trading.

Optiver’s software engineering function was profiled by The Pragmatic Engineer, a leading technology publication on Substack.

Latency and research iteration pull in different directions. Software engineer Daniel B. explores how three generations of trading systems evolved to balance both.

For the past several years, the AI conversation has largely centered on models: which are the most capable, which will dominate and how quickly they will improve.

Engineering the Agentic SDLC

AI Engineering

Instead of throwing agents into an SDLC that wasn’t built with them in mind, we’re trying to do something different: design the development lifecycle around them from the ground up. Below is an example of how we shrank a block of work we take on when we connect to a new exchange. Over the last few months, we've been running an experiment. If you take agentic coding seriously, what does it look like?

We sat down with Pat Cooney, Head of Platform Engineering, to talk about what Platform Engineering means here and where agentic AI fits into the picture. Pat has spent over a decade at Optiver across markets, regions, and roles.

Click below

Learn more about Optiver