Trang chủEsportsBefore Talking About Winning or Losing: PPDA 25.1, 564 Minutes and the Trap of an Empty Dataset

Before Talking About Winning or Losing: PPDA 25.1, 564 Minutes and the Trap of an Empty Dataset

Câu trả lời cốt lõi: Bài viết phân tích cách nhà báo dữ liệu Đỗ Nam kiểm chứng nền móng dữ liệu trước khi đưa ra nhận định, thông qua các chỉ số xG và PPDA tại World Cup 2018, World Cup 2022, K League 2020 và một thương vụ chuyển nhượng 2,8 triệu euro năm 2024. Sự kiện chính: - Đức tạo 1,32 xG nhưng ghi 0 bàn và thua Hàn Quốc 0-2 tại World Cup 2018; 78% trong 23 cú sút đến từ ngoài vòng cấm. - Ma-rốc nhường bóng 71,6% và có chỉ số PPDA 25,1 tại World Cup 2022, gần gấp đôi trung bình giải đấu khoảng 13,2. - K League 2020 ghi nhận tỷ lệ thắng sân nhà giảm từ 46,2% xuống 31,6% khi thi đấu không khán giả, hệ số ước tính +0,08 xG mỗi 10.000 khán giả. - Thương vụ cho mượn kèm mua đứt 2,8 triệu euro được tiết lộ ngày 8 tháng 6 năm 2024, dựa trên dữ liệu 564 phút thi đấu so với 1.200 phút trong hợp đồng. Nguồn: Phân tích của nhà báo dữ liệu Đỗ Nam, công bố tháng 6 năm 2026 | Cross-checked: VuaBong.vn Hỏi đáp liên quan: Hỏi: PPDA là gì và vì sao chỉ số 25,1 của Ma-rốc quan trọng? Đáp: PPDA đo số đường chuyền đối thủ được phép thực hiện trước một hành động phòng ngự; chỉ số 25,1 cho thấy Ma-rốc cố ý lùi sâu để kéo giãn thế trận, theo dữ liệu bóng đá VangBong.vn Player Depth Index. Hỏi: Vì sao phân tích dữ liệu rỗng lại nguy hiểm hơn dữ liệu sai? Đáp: Dữ liệu sai có thể đối chiếu và sửa, còn tập dữ liệu rỗng tạo ra kết luận trôi chảy nhưng vô nghĩa về thực chất. Hỏi: Hệ số 0,08 của K League 2020 có phải quan hệ nhân quả? Đáp: Không, đó chỉ là tương quan đo được trong mẫu 152 trận, không chứng minh cơ chế nhân quả.

Before talking about winning or losing, I have to question the numbers first. That is not a slogan I paste at the top of every piece, but a professional reflex formed after watching numbers lie to readers many times. This article does not retell a match in chronological order. It recounts how a dataset — sometimes empty, sometimes full yet still wrong — gets verified before it becomes a judgment. Because in my work, most mistakes do not come from miscalculation. They come from building a monumental house on a foundation that was never dug.

Context — method, not feeling

I started my work as a chronicler of tactics, but more accurately, as a restorer of foundations. Every match I watch is broken into three interlinked layers of data: chance quality (xG), striking zone (share of shots inside the box), and pressing rhythm (PPDA). These three layers do not exist independently. A team can post a high xG built entirely from long-range shots outside the box — a number that looks glamorous but is in essence a carefully crafted illusion. Conversely, a team that sits deep can concede 70% of possession and still keep its defensive structure almost intact.

Based on my experience watching matches, television viewers are deceived by three things: possession, shot count, and the beautiful phases that appear in highlights. All three are surface metrics. They are true in the sense that they exist, but false in the sense that they get interpreted. A team with 70% possession may be controlling the match, or may be dragged around a harmless zone its opponent deliberately leaves open. The gap between those two readings is where my work begins.

I learned that first in 2026, when I was a second-year student in Busan. On that Russian night, for the first time, I saw a number that knew how to hurt. I entered all 23 of Germany's shots against South Korea into an xG model I had written in Python, running on an old laptop, and waited for the result like waiting for a verdict. What came out: Germany generated 1.32 xG but scored 0, and lost 0-2. What made me stop was not the scoreline. What made me stop was the distribution of those shots.

Before Talking About Winning or Losing: PPDA 25.1, 564 Minutes and the Trap of an Empty Dataset

I cross-checked the model against the highlights and realized the naked eye was being systematically fooled by the feel of the ball. Of those 23 shots, 18 — 78% — came from outside the box. In other words, the reigning world champions poured the ball toward the opponent's goal in a way that looked utterly dominant, but most of the effort ended where the probability of becoming a goal fell below any meaningful threshold. That was not a miracle. It was the consequence of a tactically unwise decision, repeated until it became habit, until it became a sentence.

Since that night, I have set a personal rule I have never broken: never write a judgment based only on highlights or a commentator's words. Every claim must be anchored to a structure of at least three numbers, accompanied by a question few bother to ask: where does this data come from, how many matches is the sample, and who selected it.

Four markers, one method

Four markers shaped the way I read sport to this day. I tell them not strictly in chronological order, but in the order in which they changed me.

The first marker is 2026, when I began my career as an esports athlete and tournament organizer before moving into esports media. That period taught me that every competition, whether on a pitch or in a digital arena, runs on a set of implicit rules. Outsiders see the result; insiders see the structure that produces it. I left the athlete's chair but kept the habit: always ask about the rules and conditions before asking about the winner.

The second marker is 2026 — the Russian night, the night a number first knew how to hurt in my hands. It taught me that surface impressions and data truth can diverge so far that a champion looks dominant while it is actually lost.

The third marker is 2026, when K League 1 became one of the first top leagues in the world to resume in front of empty stands. The xG model I built in 2026 began to drift. Not the kind of drift you can wave away as a small margin of error, but the kind where a foundational variable has vanished. I decided not to fix the conclusion but to fix the assumption. I collected 152 matches and found the home win rate fell from 46.2% in the 2026 season to 31.6% under no-spectator conditions. I completed a 40-page report concluding that every 10,000 spectators were equivalent to about +0.08 expected goals for the home team.

The 0.08 coefficient did not measure the silence; it measured what we had lost. Nobody asked for that report. But I knew that if I did not fix the foundation, every subsequent analysis would be wrong — wrong not at the layer of judgment, but at the layer of assumption.

The fourth marker is December 2026, in Morocco. That is where I learned to use data to defend a position against the crowd.

The PPDA 25.1 defense and the lesson of choosing to sit deep

In December 2026, I was assigned to analyze Morocco — the first African team to reach a World Cup semifinal. The media at the time, especially the Korean media where I was working, described Morocco with very familiar phrases: resilient, gritty, parking the bus, lucky. Those are emotional words draped over a phenomenon that could be fully explained by numbers.

I compiled data from Morocco's knockout run and rebuilt the story. That team conceded 71.6% of possession on average. On the surface, that is the hallmark of a side suffocated by pressure. But when you place that figure beside PPDA — the metric for how many passes a team allows an opponent before making a defensive action — the picture flips entirely.

The most shocking figure was PPDA 25.1, nearly double the tournament average at the time of about 13.2. In other words, Morocco let opponents pass many times before engaging in a real challenge. But that was not passivity. It was a deliberate choice: to let opponents pass in zones where passing creates no danger.

PPDA 25.1 — sitting deep is not a concession, it is stretching the shape. I wrote that line as a way to seal the whole argument. Morocco's side accepted letting opponents hold the ball, pulled them out of their attacking structure, drained their time and space, then punished them with lightning counters or with a back line organized so tightly that high-quality chances barely existed.

Before Talking About Winning or Losing: PPDA 25.1, 564 Minutes and the Trap of an Empty Dataset

Across that run, the opponents' total xG against Morocco reached 4.02, yet the actual goals conceded were significantly lower. That is proof of something xG models sometimes have to bow to: not every chance becomes a goal, and not every defense is fairly judged by counting goals conceded.

Every shot that hits the post is a world never born. When an xG model assigns a chance a value of 0.3 goals, it is not saying the chance will score 3 out of 10 times. It is saying that in a large sample, similar chances scored at a frequency of about 30%. In a single match, 0.3 is a promise not guaranteed. And the gap between promise and reality is where many analysts sink. They believe the promise instead of checking the promiser.

My Morocco piece argued a counter-current point: defending does not mean being passive. From then on I replaced the phrase "pinned back" with "choosing to sit deep" whenever describing an organized defensive side. Not because the new phrase sounds better. Because the old phrase misdescribed the nature of the phenomenon, and a misdescription breeds wrong analyses at the next layer.

The chain of evidence — when the crowd loves a wrong story

The most interesting thing about Morocco was not how far they went. The interesting thing was how they were misread for nearly a month, even while public data was sufficient to point out the truth.

When a team concedes 71.6% possession, the viewer's reflex is to think of weakness. But holding the ball has never been proof of control if most of that possession happens in midfield and the two flanks, where a misplaced pass merely lets the defending side recover and reset its structure.

Across three knockout matches, the total xG Morocco conceded came to roughly 4.02 combined. That figure, given the context of knockout games against elite attacking sides, is low enough that statisticians must recheck their own data. A back line "suffocated" by pressure should have faced a far higher total xG. The fact that it stayed at 4.02 is evidence that most of the possession opponents enjoyed never converted into quality chances.

Here a familiar blind spot of sports media appears. We like telling stories about heroes. But we also like telling stories about underdogs rising through character. Both narratives are compelling, and both easily push us to skip verifying the foundation. The right question is not "does Morocco have character?" but "how did Morocco's defensive structure neutralize the opponent's attacking structure?"

To answer that, I had to step back further than the match itself. I had to distinguish three things: the zones where opponents held the ball, the value of passes in those zones, and the moments Morocco chose to press. When those three layers were stitched together, a model emerged as clear as a map: Morocco gave opponents freedom in non-dangerous zones, kept the distances between lines almost constant, and only increased pressure when the ball moved into zones that could yield a real chance. That was not luck. That was a system.

And when a system works, the scoreline is no longer the only proof of quality. It is merely the final result of a chain of decisions, including some poor decisions by others.

The contrarian angle — correlation is not causation

There is one mistake data people are most prone to, and it is also the mistake readers are most easily fooled by: mistaking correlation for causation.

When I published the 0.08 coefficient for K League 2026 — every 10,000 spectators equal to about +0.08 expected goals for the home team — the first reaction of many was to turn it into an absolute truth. They said: so whenever there are spectators, the home team gets stronger. But the data does not say that. The data says only that in that specific sample of 152 matches, under those specific conditions, a measurable correlation exists. It does not prove the causal mechanism, and it does not guarantee the correlation will repeat in another season, another league, or another country.

This is where an analyst must be honest with himself about the model's limits. A number strong enough to look like truth is often precisely the number that needs the most scrutiny. Whenever I publish a coefficient, I try to include three things: the estimated error, the sample size, and the assumptions the model must accept to operate. Remove one of the three and I have turned analysis into propaganda.

I believe data analysts are entering the dressing room in a metaphorical sense, and not always in a good sense. Their conclusions are often detached from the actual rhythm of the match — the player's body, the team's psychology, the things happening over seconds no model captures. Data can show that a team should change its approach. But data cannot play for the players, and cannot sit on the bench during a tense period of extra time.

That is why I always keep a silence between the number and the conclusion. That silence is not hesitation. It is respect for the reality data has not yet measured.

2026 — 564 minutes and the Lisbon connection

In 2026, when I was 25, my Morocco piece brought me a connection I had not expected: a sports data company based in Lisbon. From that source, I gained access to a different kind of data than I was used to — not on-pitch xG, but players' match logs, minute by minute, entry by entry, substitution by substitution.

And I discovered a case that forced me to rewrite how I read transfer news. A Korean midfielder at a mid-table club had played only 564 minutes the previous season — far below the 1,200 minutes recorded in his contract. The gap between contract and actual playing time is a signal. Not a single signal, but a signal within a chain.

I sent the player's agent a six-page metric report in which I did not judge whether he was good or bad, but simply presented the facts: minutes played, number of substitute appearances, average minutes per appearance, and the gap between contract expectations and on-pitch reality. On June 8, 2026, I was the first to reveal the deal — a loan with a 2.8 million euro purchase option.

Transfer fees do not measure talent; they measure the buyer's desire. The 2.8 million euro figure does not say how good the player is. It says that a club, at a specific moment, was willing to pay that much for a belief. And that belief, in turn, must be verified by data rather than inspiration.

The agent later shared that they trusted me because I brought numerical evidence, not emotional judgment. That is the compliment I treasure most in my career, not because it praises me, but because it confirms my method was on the right track even in a field ruled by emotion.

Since then my transfer reporting has had a fixed logical frame: hypothesis, data, source, probability. I dropped vague terms like "declining form" and replaced them with verifiable statements: "minutes played fell 41% versus the previous season." The difference between the two ways of writing is not style. It is that one can be proven wrong, and the other cannot.

The Russian night, and the trap of an empty dataset

There is a lesson I had never written down until now, though it has haunted me for years. It is the lesson of an empty dataset.

In analysis, people fear wrong numbers most. But in my experience, what is more dangerous is numbers that do not exist yet are treated as if they do. An analysis built on wrong data can still be checked and corrected, because it has something to compare against. But an analysis built on an empty dataset — or on a sample too small to represent anything — will flow like a fine essay, every sentence formally correct and substantively meaningless.

When my xG model drifted in 2026, the first thing I did was not look for calculation errors. It was to find the variable that had vanished — spectators. Had I skipped that step and kept running the model as before, all my subsequent conclusions would have carried a systematic error, one that does not appear at the results layer but hides at the assumption layer.

This is exactly the point where many sports analysts, especially newcomers, stumble. They jump from one match, one phase, one striking number to a sweeping conclusion. They say a team has declined after one loss. They say a player is finished after three missed shots. They build long-term claims on a foundation of a sample of one. Had I ever built judgments that way, I would have destroyed my own credibility long ago.

Every meta update is a confession by the publisher. In esports, this holds in the sense that a patch changing the rules often admits the previous version created an unwanted state — an overpowered champion, an over-advantaged tactic, an overly dry gameplay loop. But the line also holds in football metaphorically: whenever foundational conditions change — spectators vanish, the schedule tightens, offside rules adjust — old numbers quietly lose value, even as they still appear on screen with the same trustworthy face.

And when that happens, readers have no way to know they are being led by dead numbers. They only see a rigorous analysis, with figures, citations, clear conclusions. The death lies in the foundation no one sees, because we never look down at the ground while being carried away by a good story.

The writer's blind spot and the reader's blind spot

There is a paradox in the work of a tactics chronicler who works with data: the deeper you analyze, the more easily you create new blind spots.

The first blind spot is the habit of treating a strong metric as an absolute truth. When I spend great effort verifying a coefficient, I tend to assign it too much weight, because I have invested in it with both time and reputation. This is a purely psychological trap: people defend what they have built, even when new evidence says it needs adjusting.

The second blind spot is sitting deep to excess. I have a habit of weighing every angle before reaching a conclusion, and that habit, if unchecked, can turn an analysis into an endless string of conditions that never commits to anything. Smart readers need caution, but they also need a conclusion to weigh. A piece that ends with ten hypotheses and no judgment is a piece dodging responsibility.

The third blind spot is a lecturing tone. When you spend years fighting information noise and crowd sentimentality, you easily develop a tone that looks down on the reader. That is what I fear most in myself. The right assumption is not that the reader is inferior. It is that the reader is intelligent but has not yet formed the habit of reading numbers. My job is to build that habit, not to judge its absence.

The fourth blind spot, and perhaps the most important, is confusing accuracy with completeness. A dataset can be perfectly accurate yet incomplete, and an analysis based on it will still be wrong. This is what data people rarely admit: most big mistakes do not come from wrong numbers, but from right numbers in a missing set.

The Russian night, reread years later

Every time I reread the Russian night analysis, I see something new. In 2026, I saw an attacking system with a skewed distribution. In 2026, I saw a dataset whose foundational conditions had changed. In 2026, I saw a lesson in how data can defend a counter-current position. In 2026, I saw a lesson in how a number can become a transfer signal.

But by 2026, rereading it, I see the most important thing is not the 1.32 xG figure, nor the 78% of shots from outside the box. The most important thing is that I dared to say the reigning world champions were eliminated not by a miracle, but by a series of tactically unwise decisions repeated until they were inevitable. Saying that at the time was not popular. It went against the expectations of the crowd, who wanted to believe in a historic event as a fairy tale.

And I believe that, precisely at that point, my work became valuable. Not because I predicted a match correctly, but because I refused to explain a phenomenon with something that could not be verified.

The transfer market and the price of belief

The transfer race among the giants always draws my attention more than it truly deserves. Read the news and you will see a flood of figures: 100 million, 150 million, 200 million euros. These numbers are shocking enough to become headlines, but they often measure something other than talent.

From my 2026 experience with the 2.8 million euro deal, I drew one observation: the most analytically interesting contracts are not always the biggest. They are often contracts at small clubs, where every euro must be weighed carefully, where the data profile decides more than the buyer's brand.

At a big club, a transfer fee reflects part talent and part brand ambition. At a small club, a transfer fee often reflects a hypothesis about the future — the hypothesis that a player who played only 564 minutes can play 2,500 minutes if placed in a suitable system. That is the kind of transaction worth analyzing, because it rests on a chain of reasoning that can be verified or refuted.

The problem with today's transfer market is that it is gradually losing that kind of reasoning. The numbers are big enough that they are themselves a story, and once a number is a story, it no longer needs analysis. This is where I always try to push back in my writing: treat a 100 million euro deal with the same rigor as a 2.8 million euro deal, and never let the scale of a number overshadow the quality of the evidence.

When data enters the dressing room

There is a truth I constantly remind myself of: data cannot replace people in the dressing room.

This sounds obvious, but it is violated constantly. As analytical models grow more sophisticated, the pressure on coaches to please the numbers grows too. A coach can be criticized for picking a player with low metrics but suited to his system. A player can be judged by a metric that does not reflect his true role on the pitch.

This is what I call the detachment between data conclusions and actual rhythm. A model can say Player X should start every match. But the model does not know that Player X is dealing with personal problems, or that he has just moved to a new city, or that the system the model rates him within no longer exists since the coach changed the formation.

I am not against data entering the dressing room. I am against entering without knocking. In my writing, I always try to leave space for what the model cannot measure, and I always make clear that a metric is never a final verdict.

Moving up from the foundation

Looking back on the journey from 2026 to now, I realize the single principle that has never changed in my work is verifying the foundation before building the judgment layer.

It means that before saying a team won because of character, I must check whether its tactical structure created a measurable advantage. Before saying a player is finished, I must check minutes played, positions taken, and the quality of the chances he was given. Before saying some coefficient predicts a match outcome, I must ask about sample size, about error, and about the assumptions the model is forced to accept.

This principle does not make my work faster. It makes it slower. Each analysis costs me many hours just to gather background data, cross-check sources, and discard numbers that cannot be verified. But that is a price I am willing to pay, because I know that once the foundation is wrong, everything built on it must be torn down and rebuilt.

I do not write about football. I write about the light that data casts. And that light only has value when it is aimed at a foundation solid enough to withstand it. When the light shines on an empty dataset, it does not reveal truth. It only reveals emptiness, and how we filled it with belief instead of evidence.

Signal for the next round

If I have to place one judgment for the next round, I will not place it on any team or any specific match. I will place it on the quality of the datasets we use.

As automated data platforms expand, as models grow more complex, the biggest risk is not that we will calculate wrong. The biggest risk is that we will analyze, very rigorously, an empty dataset — numbers that represent nothing, presented with the confidence of a well-designed model.

Every shot that hits the post is a world never born, and every empty dataset is a world that never existed. The chronicler's job is to tell the two apart before writing anything. Because a match can end and be forgotten, but a model running on a false foundation will quietly steer thousands of subsequent judgments, and among them are judgments a team must live with for an entire season.

Before Talking About Winning or Losing: PPDA 25.1, 564 Minutes and the Trap of an Empty Dataset

The question I leave this time is not which team will win it all. The question is: next time you read an analysis that opens with an impressive number, will you pause long enough to ask where that number comes from, how many matches the sample holds, and what has been left out?

If the answer is yes, then the foundation has begun to be dug.

Cầu thủ liên quan