🇮🇷 TRACK VESSEL ACTIVITY IN THE STRAIT OF HORMUZ 🇮🇷

MONITOR LIVE

GUIDES

Verification Is the New Visibility: Why Data Quality Decides What Maritime AI™ Is Worth

What’s inside?

    Executive Summary

    • Maritime data is not messy by accident but manipulated on purpose, by operators with a commercial or strategic reason to make a vessel appear somewhere it is not, owned by someone who does not control it, or doing something it never did.
    • Since data can be manipulated, a single source is not always enough to confirm a record is correct, and corroboration is what makes it reliable.
    • AI raises the cost of unverified data rather than lowering it, removing the manual review step that used to catch implausible positions and propagating errors into every downstream product at machine speed.
    • Manufactured risk costs as much as missed risk, with GPS jamming inventing port calls, meetings, and dark activity that inflate a vessel’s risk classification while its behavior never changed.
    • Verification is an operating discipline rather than a data-cleaning project, running daily across 150,000+ vessels, keeping a human on the cases where judgment changes the answer, and feeding every confirmed outcome back into the model that raised the case.

    Some Maritime Data Is Wrong on Purpose

    Hundreds of vessels carry a lower compliance risk classification today than they did before their data was reviewed. Most moved from high risk to low. Their behavior never changed. What changed is that port calls, meetings, and dark periods recorded against them were confirmed as artifacts of GPS jamming and removed.

    Those vessels were never high risk. For as long as the false activity sat on their records, every system reading those records believed they were, and every decision taken on that basis inherited the error.

    This is the condition that separates maritime data from almost every other data problem an enterprise or an agency deals with. Most data quality work assumes errors are accidents. A field is missing, a system times out, two databases drift apart. Maritime data contains all of that, and then a second category on top of it. Positions are broadcast that never happened. Identities are split, merged, and reissued. Ownership is layered through structures built specifically to survive scrutiny. Signals are degraded across entire regions by actors with the means and the motive to degrade them.

    The distinction matters operationally. Accidental error is random, which means it tends to cancel out at scale and rarely points in a convenient direction. Deliberate error is adversarial. It is designed to survive the checks a reasonable analyst would run, it concentrates precisely where the stakes are highest, and it always points toward a conclusion someone wants a decision-maker to reach.

    Counting Sources Is Not the Same as Checking Them

    Ask a maritime data provider about quality and the answer usually arrives in units of quantity. Number of sources, number of vessels covered, update frequency, percentage of the global fleet visible. Those describe coverage, and coverage is worth having. None of them describes whether a given record is true.

    The gap between the two is where operational failure lives. A provider can ingest more sources than anyone in the market and still publish a position that never happened, because ingesting a signal is not the same as validating it. A provider can update every record every minute and propagate an error faster than a slower one would. Coverage and correctness are different properties, and holding more of the first does not deliver the second.

    The underlying assumption is that aggregation resolves conflict. Gather enough sources and the truth emerges from consensus. That assumption holds when sources fail independently, and sensors do not fail in the same way as each other. A manipulated transmission and an inconclusive image are different problems, which is why corroboration depends on knowing how each source fails and which sources can settle the question, rather than on how many are held.

    This is what all-source operational intelligence is for. Windward fuses AIS, EO and SAR satellite imagery, RF emissions, digital presence signals from onboard devices, non-AIS satellite networks, open-source material including registries and filings, and more, all resolved against a single vessel identity. The list is deliberately open rather than fixed, which is the distinction between all-source and multi-source. A platform built around the roster it launched with treats every new source as an engineering project, and the sources that will matter in five years are not necessarily the ones available today.

    All-source does not mean everything, every time. Each mission has a different question at its center, and the sources that answer it differ accordingly. Sanctions work turns on ownership and transfer behavior, force protection on what is moving in an area now, illegal fishing on vessels that never transmitted at all. What matters is that the source capable of settling a given question is available when that question gets asked, and that the platform can take in any form of data and resolve it against the same vessel identity, so a source that did not exist last year is a configuration exercise rather than a rebuild.

    The question a buyer should be asking is not how much data a provider holds — it is what the provider does when two of its sources disagree.

    AI Raises the Cost of Bad Data Rather Than Lowering It

    The expectation running through most conversations about AI in this domain is that models will compensate for imperfect inputs. Enough scale, enough pattern recognition, and the noise averages out.

    The opposite is closer to true, for three reasons.

    Automation removes the last human filter. Manual analysis was slow, and it did not scale, but it had one property worth preserving. Somewhere in the chain, a person looked at an implausible position and hesitated. That hesitation was an unlogged, unmeasured quality control layer, and automation removes it by design. An automated pipeline does not hesitate. It computes.

    Errors propagate further and faster. A bad position in a manual workflow affected one analyst’s assessment. A bad position in an automated workflow flows into the activity record, the risk score, the alert, the lead, the audit trail, and the report that goes to a regulator or a commander. The error is now embedded in six artifacts, and only one of them is traceable back to its origin.

    Explainability makes a wrong answer more persuasive. This is the least discussed and the most dangerous. A model that shows its reasoning is more trustworthy than one that does not, which means an explanation built on a fabricated port call is more likely to be believed than an unexplained score would be. Confidence is a function of the reasoning being visible, not of the inputs being true.

    The reasonable conclusion is not that AI should be held back until the data is perfect. The data will never be perfect, because an adversary is actively working to keep it imperfect. The conclusion is that verification has to be built into the same system that does the automating. Fusion on top of unverified inputs produces an assessment that is fast, confident, fully explained, and still wrong. That is more dangerous than an uncertain one, because an assessment that looks reliable is acted on without being checked.

    Every Risk Score Inherits Five Layers of Assumption

    A maritime risk assessment looks like a single output. It is actually the top of a stack, and each layer takes the layer below it on trust. An error introduced low in the stack does not stay there. It is inherited upward, gaining apparent authority at every step, until it emerges as a number that looks like a finding.

    Every risk score inherits five layers of assumption, Windward

    Read upward from the bottom, and the dependency is obvious. Read downward from the top, and it is invisible, which is how bad risk scores survive review.

    Vessel Details

    Deadweight, vessel type, length, and the other fields that determine how a vessel is classified and screened. Everything above inherits them, and they are being brought up to the same standard as ownership.

    Identity

    Two records that are really one vessel, or one voyage split across both. The quietest failure in the stack and the one that distorts the most, because a split record means a vessel’s history is only half present wherever it is read. Windward’s platform now holds a single clean record of truth for more than 95% of vessels, up from 88% at the end of 2025.

    Ownership

    Who owns a vessel, who operates it, and who ultimately controls it. Ownership is the layer most often engineered to deceive, because it is the layer that determines sanctions exposure. It is also the layer most often taken from a single feed and never challenged.

    Position

    Where a vessel actually is, as opposed to where it reports itself to be. A GNSS manipulation event is a position that was broadcast but never happened, and everything downstream inherits the inaccuracy.

    Activity

    What a vessel actually did. Port calls, ship-to-ship transfers, meetings, dark periods. This layer is derived from position, which means it inherits every positional error automatically and converts it into something that reads like observed behavior.

    Risk

    The output. It inherits all five layers beneath it and presents them as a single assessment, which is precisely why an unverified stack is so dangerous. The output format gives no indication of which layer it is standing on.

    The practical consequence is that data quality cannot be audited at the top. Checking whether a risk score looks reasonable tells you nothing about whether the activity under it happened. Verification has to be applied at each layer independently, and it has to be applied continuously, because an adversary who understands the stack will attack whichever layer is checked least often.

    Manufactured Risk Costs as Much as Missed Risk

    The failure everyone plans for is the miss. A sanctioned vessel clears screening, a dark transfer goes undetected, a threat is not seen until after the event. It is the failure that generates enforcement action and board-level consequence, and it is the reason the industry exists.

    The opposite failure attracts almost no attention and costs nearly as much.

    GPS jamming does not only hide vessels. It manufactures activity that never happened. A jammed position produces a port call the vessel never made, a meeting it never held, a dark period it never had. Each of those is a legitimate risk indicator when it is real. None of them is distinguishable from the real thing at the point it enters the record.

    A very large crude carrier waiting to load off Qatar appears in six different areas within 24 hours, the result of area-level GPS jamming corrupting the position data its AIS transmissions carry. Each false position can generate a port call or meeting the vessel never had, inflating its risk score. Source: Windward Maritime AI™ Platform.
    A very large crude carrier waiting to load off Qatar appears in six different areas within 24 hours, the result of area-level GPS jamming corrupting the position data its AIS transmissions carry. Each false position can generate a port call or meeting the vessel never had, inflating its risk score. Source: Windward Maritime AI™ Platform.

    The result is a vessel whose risk classification rises without its behavior changing. Where jamming is persistent, Windward reviews the affected area and removes the false activity across it. Hundreds of vessels now carry a lower compliance risk classification than they otherwise would, most of them moving from high to low.

    That is a measured outcome checked against daily risk snapshots rather than a reconstructed history. It is worth sitting with what it implies about the period before the review. For every one of those vessels, someone was making decisions against a risk profile that the vessel had not earned.

    For a commercial team, a false high-risk classification means a counterparty rejected without cause, a fixture delayed, a transaction escalated into enhanced due diligence that consumes days and produces nothing. For a government team, it means finite analytic and collection capacity spent on a vessel that was never the problem, which is capacity not spent on one that was.

    Both failures come from the same root, and there is no configuration that trades one against the other. Tightening thresholds to catch more real risk manufactures more false risk. Loosening them to reduce noise lets real risk through. The only move that improves both sides at once is verifying the underlying activity, which is why the jamming work removes false activity outright rather than down-weighting it.

    Windward does this at three points. Before the activity is created, where a model has prevented GPS jamming from generating false ship-to-ship meetings since May 2026. As it happens, near-real-time detection flags jammed areas so a compliance team can see that a risk decision may be affected and override it. And after the fact, persistent jamming areas are reviewed and the false activity in them removed at once.

    The governing principle is stated plainly in Windward’s operating posture. A warning we cannot stand behind is worse than no warning. About half of suspected location spoofing events do not hold up when reviewed, and those are removed rather than left on the vessel.

    What Verification Looks Like When It Actually Runs

    The distance between claiming data quality and operating it is a process that runs every day against cases nobody asked for.

    Windward runs one pipeline across every data domain. Ownership, location spoofing, GPS jamming, and vessel movement each have their own detection model, and all of them resolve through the same four steps.

    Automation carries the volume. Judgement carries the ambiguity, Windward

    Only step 1 differs by domain, built for what can go wrong in that domain specifically. Steps 2, 3, and 4 never change.

    Step 1 is not a rule engine. Language models read the unstructured material where change actually surfaces, including news, registries, corporate filings, and external sources, and pick up what never arrives as a clean feed. A company quietly changing hands. A name appearing where it should not. Analytical models watch the data itself for behavior that does not fit, whether that is a vessel moving in a way its own history does not support, or a position contradicting every other source held.

    Between them, they raise the case before anyone goes looking for it, score it, and set its priority. Where the model is confident, it applies the fix, and the case closes with no person involved. That is most of what is raised.

    The rest is the part that determines whether any of this works. These are the cases where confidence is too low to act alone. They arrive with priority already attached, so sanctioned, high-risk, and dark fleet vessels are worked first. The reviewer investigates against everything else held on that vessel, including other position sources, the vessel’s own history, and satellite imagery where the signal alone cannot settle it. When the call is clear, they make it. When it is not, a senior reviewer makes the final call.

    Two rules govern that path. Nothing ambiguous is resolved by guessing, and nothing is deleted without a human confirming it.

    This is what “human judgment is decisive” means operationally. Not a person reviewing everything, which does not scale and never did. A person placed exactly where judgment changes the answer, and automation carrying everything where it does not.

    What gets checked, and what changed this year, Windward

    The closing step is what compounds. Every confirmed outcome is fed back into the model that raised the case. Verified data is not only a cleaner input. It is training signal, which means a provider without a human verification loop has no mechanism for its models to improve against deception that is itself adapting. That loop is why turnaround and accuracy have both continued improving through the year rather than leveling off.

    The operating tempo behind it is worth naming, because it is where most quality claims quietly fail. Ownership is reviewed across more than 150,000 vessels, not only tankers and cargo, and that full review runs weekly rather than as an annual sweep. Anything it finds is corrected the day it is found. A vessel that appears twice today is typically found and merged within a day, against a median closer to twelve days a year ago. Newly verified data appears in the platform within one to two days of confirmation.

    For issues clients report, full resolution now averages 2.4 days against 6.7 at the start of 2026. Issues needing a second fix fell from 9% to 5%, and 40% fewer issues reach clients at all, because they are caught first.

    Capacity is finite, and the honest version of this includes where it does not yet reach. GPS jamming review is prioritized where jamming is heaviest, and by the cases clients raise. The Iranian coast and Russian Black Sea ports are under continuous review today. The Gulf of Oman, the Arabian Gulf, the Baltic Sea, and the Kerch Strait are being brought into the same process as reviewer capacity allows.

    For Government, Unverified Data Is a Capacity Problem

    Government maritime teams do not fail because they lack data. They fail because decision windows close faster than fragmented data can be reconciled, and because the capacity available to reconcile it is fixed.

    That makes data quality a resourcing question before it is a technology question. Every hour spent resolving a contradiction between two sources is an hour not spent on the assessment that contradiction was supposed to support. Every false activity in the picture is analytic attention allocated against something that did not happen. In a domain where collection and analytic capacity are the binding constraints, manufactured risk is a direct tax on operational tempo.

    The second consequence is defensibility. Enforcement decisions are challenged, and the challenge is rarely about whether a system flagged a vessel. It is about what the flag rested on. An assessment built on a position that cannot be corroborated against an independent source is not an assessment that survives scrutiny, whether that scrutiny comes from a court, an interagency partner, or a coalition ally working from a different intelligence picture. A common operating picture is only common if the participants trust the same underlying record.

    The third is the one that does the most damage over time. A false flag that is never resolved does not stay neutral. It trains the people who see it. Analysts who repeatedly work leads that dissolve on investigation learn to discount the system that generated them, and that learned discounting applies to the true flags as well as the false ones. Detection capability that is not trusted is detection capability that is not used.

    This is where the GPS jamming regions listed above matter concretely. They are among the most operationally consequential waters in the world, and they are exactly the areas where positional data is least reliable. Signal degradation concentrates where the interest concentrates. Any picture built on cooperative signals alone will be weakest precisely where it is needed most.

    The transition from awareness to action depends on the reliability of what the awareness rests on. Compressing the decision cycle is only an advantage if the decision is correct. A faster path to a wrong call is not a force multiplier.

    For Commercial Teams, Both Failure Modes Hit the P&L

    Commercial exposure to data quality runs in both directions, and most organizations are only defended against one of them.

    The defended direction is the miss. A sanctioned counterparty clears screening, a cargo of undeclared origin enters the supply chain, a vessel’s ownership resolves to a designated entity three layers up. The consequences are penalties, secondary sanctions exposure, reputational damage, and the loss of banking or insurance relationships. Compliance functions are built around preventing exactly this.

    The undefended direction is friction, and it is measured in deals rather than fines. A vessel flagged high risk on manufactured activity is a fixture delayed while the flag is investigated, a counterparty declined without cause, a transaction escalated into enhanced due diligence that consumes days and resolves to nothing. None of that appears in a compliance report as a failure, because a false positive that is correctly investigated looks like the process working. The cost lands on the commercial side of the house, where it shows up as slower clearance and business that went somewhere faster.

    For organizations whose advantage is speed of decision rather than brand protection, this is the more expensive failure. Bunker traders, mid-size tanker operators, and commodity houses close on timelines that do not accommodate a multi-day investigation of an event that never happened.

    Know Your Vessel (KYV™) exists because entity-level screening cannot close either gap. A counterparty check validates who you are dealing with on paper. It says nothing about whether the vessel carrying the cargo spent last week somewhere its transmissions deny, or whether its reported port call was an artifact of a jammed signal. The gap between reported reality and maritime reality is where both the exposure and the friction live. The near-real-time jamming flag, now rolling out, puts that context in the compliance officer’s hands, with the ability to override a decision the data no longer supports.

    The defensibility standard is the point worth carrying into a compliance committee. When a regulator or an auditor asks why a vessel was cleared, “our provider’s data showed no risk indicators” is not a defense. The defensible answer describes how the record was verified, what independent source corroborated it, and what happens when sources disagree. That answer is a property of the provider’s process, and it has to exist before the question is asked.

    This is also the trade enablement argument. Removing false positives does not weaken a compliance posture. It concentrates scrutiny on real exposure and clears the rest faster, which is what allows a team to make defensible decisions at the speed the business requires.

    Eight Questions That Separate Verification From Aggregation

    A procurement process that evaluates coverage will select for coverage. These questions test the thing that actually determines whether an output can be acted on, and they can be put to any provider in the market.

    1. When two of your sources disagree about a vessel, what happens next? The answer describes a process, or it does not exist. Automatic preference for one feed is a ranking rule, not a resolution. The useful answer names who investigates, what independent material they check against, and who decides when it stays ambiguous.
    2. How often is ownership reviewed, and across how much of the fleet? Ownership is the layer most often engineered to deceive and most often refreshed least. A review that covers only higher-risk segments leaves coastal and inland vessels outside review entirely, which is a known blind spot rather than a neutral scoping choice.
    3. What happens to a flag you can no longer stand behind? The honest answer is that it is removed. A provider that never retracts a flag is telling you its precision is unmeasured. About half of Windward’s suspected location spoofing events do not survive review, and those are deleted rather than left on the vessel.
    4. Can a vessel’s risk classification go down without its behavior changing? If the answer is no, the system cannot correct manufactured risk. Risk that only ratchets upward is not a risk model. It is an accumulator.
    5. Who confirms a deletion? Automated deletion is how real events disappear. The standard worth holding is that nothing is deleted without a person confirming it.
    6. How long from a data error being created to being resolved, measured as a median? Ask for the median and the trend, not the best case. Windward merges a vessel appearing twice within a day, down from about twelve days a year ago.
    7. What is your turnaround when we tell you a record is wrong? This is the question that reveals whether quality is operational or aspirational. Windward commits to a first analysis and a resolution path with timelines within one business day, against a five-working-day industry standard that used to be Windward’s own.
    8. Can your compliance team see when a risk decision may be based on jammed data? Without that, the team cannot tell a real warning from a false one at the moment it matters, let alone override it.

    A provider that can answer all eight is running verification. A provider that answers most of them with coverage statistics is running aggregation, and the difference will surface in the first decision that matters.

    Verification Is a Posture, Not a Project

    Data quality work is usually scoped as a cleanup with an end date. Fix the backlog, reconcile the records, close the program. That framing fails in this domain for a structural reason. The errors are being generated by people who are actively adapting, which means the backlog regenerates as fast as it is cleared and in whatever shape the last fix did not anticipate.

    What replaces the cleanup model is a pipeline that runs daily against whatever the data produces next. The direction of travel is toward catching problems earlier rather than repairing them later. Prevention is already partly live. Since May 2026, a model has stopped GPS jamming from creating false ship-to-ship meetings in the first place, and near-real-time flagging of jammed areas is now rolling out. That prevention is being extended from ship-to-ship meetings to port calls and dark activity, review of activities is becoming automatic so cases are confirmed sooner, and vessel details are being brought up to the same accuracy and update speed as ownership. Still ahead are vessel visibility earlier in the lifecycle, with a vessel card for newly registered vessels before their first signal, and the same review process extended into the remaining jamming-affected regions.

    More of the detection step becomes automated over time, with human-in-the-loop capacity spent where judgment actually changes the answer. The direction is deliberately not toward removing the human. It is toward moving the human to the cases that deserve one.

    This is what Windward’s Maritime AI™ rests on. The models are the reason the manipulation is detectable at scale, and the verified data is the reason the models can be trusted with a decision. Neither works without the other. Fusing all-source intelligence on top of unverified inputs produces confident, explainable, fast, and wrong. Fusing it on top of a record that has been checked, corrected, and fed back into the models that raised the case produces intelligence that holds up when it is challenged.

    Having more data was never the advantage. Knowing which of it is true is.

    See Verified Maritime Intelligence in Action