Own the Refinery
Africa’s AI Opportunity May Be in the Inputs
The AI economy is usually discussed as a race to build models. That framing almost guarantees Africa loses.
Frontier models require enormous amounts of compute, reliable power, specialised chips, deep technical talent and billions of dollars of patient capital. Those advantages are already concentrated in a handful of American and Chinese companies and the supply chains, cloud platforms and research pipelines built around them. Africa is unlikely to win by trying to recreate Silicon Valley’s AI stack locally.
There may be another way in.
Compute is concentrated. Data is distributed
As AI models become more capable, increasingly important inputs include information they cannot easily obtain from the public internet: proprietary workflows, specialist knowledge, underrepresented languages, real-world operating data, human-generated code, cultural material and high-quality examples against which their performance can be tested.
Africa has a lot of that. Sub-Saharan Africa alone moved more than $1.4 trillion through mobile money in 2025, roughly two-thirds of the world’s total mobile money transaction value, according to GSMA’s latest industry report. Behind that aggregate are continuously generated records of financial behaviour across markets where mobile money plays a role that conventional banking data alone cannot fully describe.
Africa has no monopoly on this opportunity. India, Brazil and other emerging markets have large technology workforces, substantial local datasets and established data-services industries already positioned to serve many of the same buyers. A model builder indifferent to which market supplies the relevant signal will buy from whoever offers the best combination of quality, price and speed. What resists substitution is the specific texture of a market: the way Lagos traders code-switch between languages, the rhythm of a Kenyan chama savings cycle, the particular shape of a naira devaluation working through a small business’s books. Generic “emerging-market” data can become a commodity. Precisely located operating knowledge is harder to substitute.
The internet gave AI its first corpus
The first generation of large language models benefited from an extraordinary historical accident. Humanity had spent decades putting books, websites, software, discussions, images and other knowledge online. Model developers could ingest enormous amounts of it at relatively low marginal cost.
The next frontier is harder. AI systems increasingly need to reason reliably, perform specialised tasks and operate inside particular professions, industries, languages and real-world environments. More generic internet data does not necessarily solve those problems.
A model may understand the textbook principles of credit and still struggle with lending behaviour in an economy experiencing severe currency volatility. It may speak English and still misunderstand Nigerian English. It may understand the engineering of an electricity network while having little exposure to how businesses actually operate around unreliable supply. It may generate something labelled “Afrobeats” without understanding the musical, linguistic and cultural structures that distinguish it.
These are information gaps, and information gaps can become economic assets.
Many of the characteristics that make African markets difficult to operate in also make them informationally unusual: informal commerce, multilingual communication, mobile-money-heavy financial systems, volatile currencies, legacy technology, intermittent infrastructure, distinctive cultural systems and regulatory environments that do not map neatly onto Western equivalents.
A frontier AI company can buy more compute. It cannot instantly manufacture twenty years of operating experience in an African market. It cannot recreate genuine Yoruba conversations by translating English. Nor can it synthetically invent the historical record of how businesses adapted to infrastructure failures, currency shocks or informal distribution systems and assume that the result represents reality.
This becomes more important as AI moves from knowing about the world to acting within it. An AI system performing work in Lagos, Nairobi or Accra eventually has to understand the realities of Lagos, Nairobi and Accra.
Some of the information generated by those environments may therefore acquire economic value precisely because it is difficult to reproduce elsewhere.
But possessing valuable information and capturing its economic value are very different things.
Africa has seen this movie before.
The worst outcome would be another raw-material economy
Across many of its most important global value chains, Africa has followed a familiar pattern: extract something valuable, export it in relatively raw form, allow somebody elsewhere to refine it, manufacture around it, control distribution and own the customer, then import the higher-value product.
Oil is the obvious example. Cocoa-producing economies capture only part of the value ultimately embedded in branded chocolate. Mineral-rich countries export ores while much greater value can accumulate in processing, battery materials, components and finished technologies.
The recurring problem sits in where value-chain ownership, transformation and bargaining power land, not in whether the raw material was valuable to begin with.
AI could reproduce that structure with information. African companies provide the data. African workers label it. Foreign intermediaries aggregate and structure it. Technology companies convert it into better models. Those models are then sold globally, including back into African markets.
The commodity has changed. The economics of extraction have not.
If that becomes Africa’s principal role in AI, we will have digitised the commodity-export model rather than escaped it.
That is why simply declaring that “African data is valuable” misses the point. Data is only the beginning of the value chain.
From raw data to AI infrastructure
Raw data is frequently messy, legally uncertain, technically unusable and commercially difficult to evaluate. Where the underlying information is personal rather than institutional (payments, health or behavioural data, for example), privacy and data-protection regimes also constrain how it can be processed, transferred and commercialised. The rights question therefore extends beyond who possesses the data to what they are legally permitted to do with it. Turning raw information into an AI asset looks something like:
Raw Information → Rights → Curation → Machine Readiness → Demonstrated Utility → AI Asset
Each stage can increase usefulness, scarcity and bargaining power. Information whose ownership cannot be established may have little commercial value. The same information with clear rights, appropriate privacy treatment, useful metadata and demonstrated relevance to a model capability is a different asset.
But the progression does not necessarily end with a cleaned dataset.
Proprietary information can potentially support benchmarks, evaluation suites, specialist reference systems, preference datasets and other tools that measure or improve how AI performs within a particular environment. A business might simply license years of information about how an industry operates. Or that information, combined with domain expertise, might be used to create a benchmark for determining whether an AI system can actually perform important tasks within that industry.
The first monetises information. The second begins converting information into AI performance infrastructure.
Music offers an intuitive example. A catalogue has historically been valued primarily around human consumption: streams, performances, synchronisation, mechanicals and other royalty-generating uses. Generative AI creates potential additional demand for rights-cleared musical information. But the opportunity need not stop at licensing songs. Music and the expertise surrounding it can potentially contribute to datasets, evaluation tools and other systems for determining whether models actually understand particular genres, languages or musical structures.
South Africa’s Lelapa AI has already built a working version of this move, on a small scale. Its Esethu Framework — developed with transcription company Way With Words and a University of Pretoria research group — gives local language communities formal governance over how their speech and text data is used, while its licensing model is designed to reinvest commercial value into further dataset creation rather than simply extracting the current one. The first release, an open-source isiXhosa speech corpus, offers an early example of the raw-data-to-AI-infrastructure move already being attempted, rather than merely theorised.
The same principle can apply to software, professional knowledge, healthcare, agriculture, finance, industrial processes and other information-rich domains.
That suggests a progression:
Raw Data → AI Inputs → AI Performance Infrastructure
The further an economy moves along that chain, the greater its potential to capture something resembling intellectual-property rents rather than commodity margins.
Follow the bottleneck
There is an important reason not to build this thesis around today’s demand for proprietary datasets: technology changes where scarcity sits.
Synthetic data may replace some human-generated datasets. Models may become more data-efficient. Information commanding a premium today may eventually become commoditised.
The opportunity does not necessarily disappear. The bottleneck moves.
If raw data becomes easier to obtain, reliable provenance becomes more valuable. If synthetic examples become abundant, trustworthy evaluation against reality becomes more important. If general models become commoditised, domain knowledge that determines whether they work inside specialised environments can become the constraint. As AI systems perform more real-world work, continuously generated feedback from that work may become more useful than historical training corpora.
The strategic objective is therefore not to identify one permanently scarce form of data. It is to understand where the information bottleneck is moving and own assets on the scarce side of it.
This also means not all African data is valuable. Most data is not inherently valuable, and rarity alone proves little. A commercially meaningful AI asset needs some combination of four characteristics: scarcity, because equivalent information is difficult to obtain elsewhere; utility, because it materially improves or tests a valuable AI capability; control, because someone possesses sufficient rights to permit the relevant use; and defensibility, because competitors cannot cheaply manufacture an adequate substitute.
Quality, freshness, longitudinal depth and domain expertise can strengthen those characteristics. Some African datasets will satisfy them. Many will not. Some will be valuable but legally unusable; others unique but irrelevant to model performance. The opportunity requires asset selection, not a continental rush to monetise databases.
And the strongest assets may ultimately not be databases at all.
A static dataset is finite. A system that continuously generates difficult-to-replicate information is different. Businesses, professional networks, creative communities and institutions produce new domain knowledge through ordinary activity every day. Their strategic asset may eventually be less the historical database than the continuing information supply chain.
That changes the economics from a one-off data sale towards something closer to recurring intellectual property.
The eventual winners may therefore not be whoever possesses the largest dataset today. They may be whoever controls systems that continually convert proprietary human activity into useful machine intelligence.
Rethinking Africa’s place in the AI stack
The conventional AI stack runs roughly:
Chips → Compute → Models → Applications
That is useful for understanding where capital is concentrated. It is incomplete for understanding where countries outside that concentration might participate.
Put another chain alongside it:
Proprietary Information → AI Inputs → AI Performance Infrastructure
Now the geography changes.
Compute rewards enormous capital investment and physical concentration. Information originates everywhere. The strategic question is whether its owners recognise its value early enough, and control enough of its transformation, to capture it.
This also complicates the increasingly popular idea of “sovereign AI.” Economic sovereignty does not necessarily require every country to build a frontier model. For many countries, that would be an extraordinarily expensive way to enter a race whose leaders already possess massive scale advantages.
Sovereignty can also mean controlling scarce assets on which other people’s systems depend.
For Africa, the objective should therefore be clear:
Do not export African data as raw material. Productise it before it leaves.
That does not mean hoarding information or pursuing digital protectionism. It means understanding rights before commercialisation, building the technical and institutional capabilities that make information more valuable, retaining ownership where appropriate, moving towards higher-value AI inputs where the economics support it, and preserving participation in the recurring value generated from continuously refreshed information.
Africa does not need to own every layer of the AI economy. It does need to become much more deliberate about which layers it allows others to own around African assets.
Technology waves create wealth across different parts of their value chains. The strategic task is to identify where a country’s actual comparative advantages sit rather than imitate the layers where others already possess overwhelming advantages.
For Africa, one candidate is becoming increasingly interesting: the proprietary information required to make artificial intelligence genuinely understand and operate across the rest of the world.
Africa has spent decades generating that information without thinking of it as an economic asset. AI may finally give some of it a price.
But price alone is not the objective. Africa already knows how the story ends when it owns the raw material while somebody else owns the refinery.
This time, the opportunity is to own more of the refinery too.


