Colonial Bias in Online Knowledge
- NOTOC**
Colonial Bias in Online Knowledge
The internet is often described as a universal storehouse of human knowledge, but the information available online reflects deep historical inequalities. Colonialism shaped which languages became dominant, whose histories were recorded, which institutions controlled archives, and what forms of knowledge were treated as authoritative. These inequalities continue in search engines, social-media platforms, digital libraries, Wikipedia, Wikidata, artificial-intelligence systems, and other technologies that organize and distribute information.
Online knowledge is not simply a neutral collection of facts. It is shaped by decisions about what is published, digitized, translated, indexed, classified, preserved, recommended, and cited. Communities whose knowledge has historically been transmitted orally, locally, collectively, or in languages with limited digital support are often poorly represented. At the same time, corporations and institutions based in the Global North possess disproportionate control over digital infrastructure, data collection, artificial-intelligence development, and the rules governing online visibility.
Colonial bias in online knowledge therefore involves more than inaccurate or offensive content. It includes the structures that determine who can produce knowledge, whose testimony is accepted, who owns the resulting data, and who benefits economically and politically from its use.
Digital Colonialism, Artificial Intelligence, and Data Extraction
Digital colonialism describes systems in which powerful corporations, governments, and institutions extract data, labor, knowledge, and economic value from less powerful populations. Although these systems do not always involve direct territorial rule, they can reproduce colonial relationships of dependency, extraction, and external control.
Large technology companies collect enormous amounts of information about people's behavior, identities, languages, movements, relationships, and preferences. This information is converted into commercial data, used to train artificial-intelligence systems, and incorporated into products controlled largely by companies headquartered in wealthy countries. The people and communities from whom the information originates may have little authority over its collection, interpretation, storage, sale, or reuse.
Artificial-intelligence systems intensify these concerns because they are trained on vast quantities of online material. The internet, however, does not contain an equal or representative record of humanity. English-language publications, Western institutions, commercially successful websites, and highly connected populations are disproportionately visible. Knowledge maintained through oral traditions, community relationships, local publications, Indigenous protocols, or low-resource languages is more likely to be absent.
As a result, AI systems may present Western categories, histories, cultural assumptions, and social norms as if they were universal. They can misrepresent Indigenous and non-Western communities, perform poorly in underrepresented languages, and reproduce stereotypes already embedded in their training material.
The AI industry also depends on workers who label data, moderate content, translate text, transcribe speech, and review disturbing material. Much of this work is performed in the Global South for comparatively low wages. The economic value produced by these workers is concentrated primarily among technology companies and investors elsewhere, creating another form of unequal extraction.
Digital colonialism also produces technological dependence. Governments, schools, researchers, businesses, and cultural institutions may rely on foreign cloud services, software platforms, identification systems, educational tools, and AI models that they cannot independently audit or govern. This dependence can weaken local control over public infrastructure and limit the ability of communities to determine how technology should serve their own priorities.
Wikipedia, Wikidata, and Knowledge Gaps
Wikipedia is one of the world's most influential sources of public information and an important source of training material for search engines and artificial-intelligence systems. Its open-editing model has greatly expanded public participation in knowledge production, but it also reflects inequalities in access, geography, language, gender, race, education, and institutional authority.
A large share of Wikipedia's contributors and content has historically been concentrated in Europe and North America. Places with fewer editors, weaker internet access, limited digitized documentation, or less coverage in conventional publications often receive fewer and shorter articles. This creates an uneven map of the world in which already prominent people, institutions, and locations become increasingly visible while less documented communities remain obscure.
Wikipedia's requirements for reliable published sources can create a particular problem for Indigenous knowledge. Colonial publishing systems frequently ignored, suppressed, or distorted Indigenous histories and oral traditions. When Wikipedia accepts only forms of documentation produced by those systems, communities can be placed in a circular dilemma: their knowledge is excluded because it was not published, but it was not published because colonial institutions did not recognize it.
Notability rules can produce similar inequalities. People from communities with less media coverage may struggle to meet standards developed around the publication practices of dominant societies. Biographies of women, Indigenous people, racial minorities, people from the Global South, and other underrepresented groups may be challenged or deleted because the historical record provides fewer conventional sources about them.
Wikidata transforms information into structured statements that can be searched and processed by machines. Although this structure is useful, databases are not culturally neutral. Decisions about categories, properties, labels, occupations, identities, locations, and relationships can reproduce assumptions inherited from colonial classification systems. Missing records and unequal coverage can also affect AI applications that use Wikidata as a source of factual knowledge.
Because Wikipedia and Wikidata are widely reused, their gaps extend far beyond Wikimedia projects. Search engines, voice assistants, educational products, automated summaries, and generative-AI systems can inherit and amplify the same imbalances. Reducing bias in AI therefore also requires improving the diversity, sourcing, governance, and geographic reach of the public knowledge systems on which AI depends.
Search Engines, Platforms, and Algorithmic Visibility
Search engines and social-media platforms determine which information is easy to find and which remains effectively invisible. Ranking and recommendation systems are designed around measurements such as popularity, engagement, advertising value, previous user behavior, and the availability of machine-readable content. These measurements tend to favor languages, institutions, and communities that already have a strong online presence.
Search results can reproduce racial, gendered, cultural, and national stereotypes. Commercial ranking systems may elevate sensational, discriminatory, or sexualized material because it attracts attention or reflects patterns already present in the data. Autocomplete systems, image searches, and recommendation algorithms can make these patterns appear objective even though they result from commercial priorities and unequal information environments.
Language is a major factor in algorithmic visibility. A platform may technically support many languages while providing its best search, moderation, translation, speech-recognition, and recommendation tools only in a small number of dominant languages. Speakers of low-resource languages may encounter fewer relevant results, more errors, weaker safety protections, and greater pressure to communicate in English or another colonial language.
Recommendation algorithms can also reinforce cultural dominance. Users seeking local-language media may be directed toward content in a more powerful regional or international language because the platform has more data, advertisers, and engagement history for that language. Over time, this can reduce the visibility and economic viability of local cultural production.
Automated content-moderation systems often perform poorly outside well-resourced languages and familiar cultural contexts. They may fail to recognize harassment, political repression, coded speech, or dangerous misinformation in marginalized languages. They may also incorrectly remove legitimate cultural or political expression because the systems lack local knowledge.
Platform ownership is therefore a knowledge issue as well as an economic issue. A small number of companies control major channels through which people search, communicate, learn, organize, and encounter public information. Their technical standards and business models influence which languages flourish, which cultural materials circulate, and which accounts of reality become most visible.
Indigenous Data Sovereignty and Knowledge Rights
Indigenous data sovereignty is the right of Indigenous peoples to govern the collection, ownership, interpretation, storage, and use of data relating to their communities, territories, cultures, languages, and environments. It challenges the assumption that information becomes freely available for outside use simply because it has been digitized or placed online.
Many conventional data practices focus on the rights of individual users. Indigenous governance frequently emphasizes collective rights and responsibilities because knowledge may belong to a family, nation, clan, ceremonial group, or community rather than to a single person. Certain materials may be restricted according to age, season, gender, kinship, spiritual responsibility, or cultural protocol.
The CARE Principles for Indigenous Data Governance emphasize Collective Benefit, Authority to Control, Responsibility, and Ethics. They complement data-management approaches that focus mainly on making information findable, accessible, interoperable, and reusable. Under a CARE-based approach, technical accessibility does not override community authority.
Free, prior, and informed consent is especially important when Indigenous languages, cultural practices, biological information, environmental observations, traditional medicines, images, recordings, or sacred knowledge are used in research or artificial-intelligence development. Consent should be meaningful, ongoing, understandable, and obtained before collection or reuse occurs.
Community authority can also be expressed through tools such as Traditional Knowledge Labels and Biocultural Labels. These labels communicate culturally specific rules about attribution, circulation, seasonal use, restricted access, and community responsibility. They help digital platforms, museums, archives, and researchers recognize that conventional copyright law may not adequately protect collective knowledge.
Indigenous data sovereignty does not necessarily require rejecting digital technology. Communities have used databases, archives, mapping systems, language tools, and digital-repatriation projects to preserve knowledge and reconnect dispersed cultural materials. The central question is whether the community controls the technology and determines how the information is described, accessed, interpreted, and shared.
Language, Translation, and Online Knowledge Inequality
Language inequality is one of the most significant sources of colonial bias online. Thousands of languages are spoken around the world, but only a small proportion have extensive digital content, high-quality translation tools, speech-recognition systems, keyboards, educational resources, or strong representation in major knowledge repositories.
English dominates much of international publishing, software development, scientific communication, and online documentation. Research written in other languages is less likely to be indexed, cited, translated, or incorporated into global policy. This can create the false impression that knowledge is absent from a region when the actual problem is that the knowledge is not available in a dominant publication language.
Scientific assessments that exclude non-English studies can overlook locally documented species, ecological changes, agricultural practices, health conditions, and social experiences. The resulting evidence base may be less accurate while continuing to portray researchers in the Global South as contributors of data rather than producers of theory and expertise.
Artificial-intelligence systems reflect the same imbalance. Models generally perform best in languages with large quantities of digitized text and well-funded technical infrastructure. In low-resource languages, they may generate incorrect information, mix languages, misunderstand cultural references, or fail to provide useful safety information.
Language digitization can support preservation and revitalization, but it can also create new risks. Companies may collect community recordings and texts to improve proprietary products without providing compensation, control, or continuing benefits to the speakers who produced the data. A language may become technically supported while the community loses authority over its digital representation.
Community-led language technology offers an alternative. Speakers can establish priorities, determine acceptable uses, control datasets, design culturally appropriate tools, and ensure that digital resources support education, cultural continuity, and intergenerational transmission. Linguistic justice requires more than translating dominant content; it requires enabling communities to create and govern knowledge in their own languages.
Digital Archives, Metadata, and Colonial Collections
Archives, museums, universities, libraries, and research institutions hold vast collections created through colonial administration, missionary activity, military occupation, scientific expeditions, anthropology, archaeology, and the removal of cultural objects. Digitization can make these collections more accessible, but digital access alone does not decolonize them.
The selection of materials for digitization determines which histories become visible. Wealthy institutions often possess the funding and technical capacity to digitize their holdings, while archives in formerly colonized countries may struggle with limited resources. This can leave institutions that acquired materials through colonial power in control of their digital representation.
Metadata also shapes interpretation. Catalog descriptions, subject headings, geographic labels, racial categories, and object names may preserve offensive or inaccurate colonial terminology. Search systems built on this metadata can continue to present colonized peoples through the perspectives of collectors, administrators, missionaries, and researchers rather than through the communities' own descriptions.
Digitized photographs and recordings raise additional ethical questions. Images created through coercion, captivity, racial classification, or unequal research relationships can be copied and circulated globally without consent. Increased visibility may reproduce the original harm, particularly when platforms strip away historical context or community restrictions.
Digital repatriation involves returning copies of cultural materials or providing communities with greater access to them. Such projects can reconnect people with languages, songs, ceremonies, family histories, and cultural records. However, digital copies should not be used as substitutes for returning physical objects when communities seek restitution.
Meaningful digital repatriation requires shared authority. Source communities should be able to determine descriptions, access conditions, cultural restrictions, interpretation, and future reuse. Decolonizing archives therefore involves changing institutional power and governance, not merely placing more colonial records online.
Decolonizing the Internet and Advancing Knowledge Justice
Decolonizing the internet means addressing the unequal relationships that determine whose knowledge is created, preserved, trusted, circulated, and monetized. It is not a metaphorical exercise or a matter of adding diverse content to systems whose underlying rules remain unchanged.
Knowledge justice begins by recognizing that marginalized communities are producers of knowledge rather than merely populations to be studied or sources from which data can be extracted. Communities should participate in defining research questions, technical standards, categories, consent processes, and measures of success.
Representation remains important. Projects that add biographies, images, local histories, and underrepresented languages to Wikimedia and other open platforms can correct major gaps. However, inclusion should not require communities to surrender authority or translate all knowledge into categories developed by dominant institutions.
Decolonial technology emphasizes plurality rather than a single universal model of digital development. Different communities may require different rules governing identity, privacy, access, ownership, preservation, attribution, and collective responsibility. Technologies should be capable of supporting these differences instead of forcing all knowledge into standardized commercial systems.
Public institutions can promote knowledge justice by funding local digital infrastructure, community archives, language technologies, open educational resources, regional research networks, and independent public-interest platforms. Governments can strengthen privacy protections, labor standards, competition policy, data-sovereignty rules, and requirements for meaningful community consultation.
Technology companies and research institutions can contribute by documenting their datasets, auditing geographic and linguistic bias, compensating data workers fairly, respecting collective consent, supporting community governance, and allowing affected populations to contest harmful classifications and automated decisions.
Individuals can participate by supporting underrepresented-language projects, improving gaps in public knowledge, questioning the apparent neutrality of search results, seeking sources from different regions, and recognizing the limits of information produced through dominant platforms.
Conclusion
Colonial bias in online knowledge is produced by overlapping inequalities in history, language, publishing, infrastructure, wealth, institutional authority, and technological power. Artificial intelligence, search engines, social media, Wikipedia, digital archives, and structured databases do not merely reflect these inequalities. They can amplify them by turning historically unequal records into automated systems that influence education, culture, public policy, and everyday decision-making.
Creating a more equitable internet requires more than increasing connectivity or adding additional data to existing platforms. It requires shifting authority toward the communities whose knowledge, languages, labor, and cultural materials are being represented. Indigenous data sovereignty, linguistic justice, community-controlled archives, fair AI labor, plural knowledge systems, and accountable platform governance all provide foundations for this work.
A decolonized digital future would not treat one region's institutions, languages, and categories as universal. It would recognize many ways of knowing, protect collective rights, distribute technological benefits more fairly, and give communities meaningful control over how their knowledge enters and moves through the digital world.
- TOC**
Digital Colonialism, AI, and Data Extraction
Why AI Systems, Which Rely on the Internet, Poorly Reflect the Diversity of Human Knowledge
| Deepak Varuvel Dennison | Le Monde | 2026-06-16
Explains how web crawling, linguistic inequality, paywalls, oral traditions, and popularity-based ranking leave AI with only a narrow slice of humanity's knowledge.
AI Is Ushering in a New Era of Colonialism
| Scott Rosenberg | Axios | 2026-06-04
Examines how Western-dominated training data and extractive data practices can flatten Indigenous and non-Western cultures in AI-generated knowledge.
Statement by the Global Digital Justice Forum at the Global Dialogue on AI Governance
| Global Digital Justice Forum | Association for Progressive Communications | 2026-05-12
Calls for AI governance that confronts knowledge colonialism, linguistic inequality, and the marginalization of Global South perspectives.
The Hidden Cost of AI: Digital Colonialism and the Global South
| WACC Global | World Association for Christian Communication | 2026-04-29
Frames AI colonialism as a continuation of older global extraction in which wealth and decision-making concentrate in the Global North.
Algorithmic Dependence and Digital Colonialism
| Samar A. Ahmed | Frontiers in Education | 2026-01-21
Proposes a framework connecting AI dependence in Global South education to data, infrastructure, epistemic, and governance colonialism.
Robots Behaving Badly: Algorithmic Colonialism and the Consequences of AI
| Bronwyn Carlson and Tamika Worrell | The Australian Sociological Association | 2026
Uses Indigenous research ethics to argue that AI systems reproduce domination, extraction, predictive harm, and colonial assumptions by design.
From Linguistic Imperialism to Algorithmic Dominance
| J. Yang | Frontiers in Psychology | 2026
Shows how language hierarchy is being reconfigured through AI systems that privilege dominant languages and reshape users' confidence and participation.
AI-Driven Media and the Reclamation of African Linguistic Identity
| K. Aiseng | SAGE Open | 2026
Studies how colonial language hierarchies embedded in algorithms affect Setswana, Tshivenda, and Xitsonga visibility in South African digital media.
Data Colonialism and Indigenous Languages in AI
| J. C. Y. Kwok | AI and Society | 2026
Reviews language-technology initiatives and warns that proprietary AI can absorb Indigenous linguistic resources without community control or consent.
Towards a Decolonial AI Sovereignty Model for Indigenous Educational Contexts
| M. F. Mbah | Cogent Education | 2026
Proposes an Indigenous-centered model for governing generative AI in education rather than importing external technological assumptions and priorities.
Reimagining the Future of Data and AI Labor in the Global South
| Brookings Authors | Brookings Institution | 2025-10-07
Highlights the invisible labor, low pay, psychological harm, and organizing efforts of Global South workers who make AI systems possible.
Digital Sovereignty and Data Colonialism: Shaping a Just Digital Order for the Global South
| Marcus Vinicius de Freitas | Policy Center for the New South | 2025-10-03
Argues that Global South countries need sovereign infrastructure, governance capacity, and regional cooperation to resist data colonialism.
Data Colonialism in the Global South: Who Owns Asia's Digital Future?
| Suhana Roy | Yale Review of International Studies | 2025-07-26
Examines how foreign platforms and data ownership structures shape Asia's digital economy and political autonomy.
The Cultural Cost of AI in Africa's Education Systems
| James Maisiri and Solomon Musonza | UNESCO | 2025-07-20
Warns that imported AI education tools may erase African values and knowledge unless local communities control design, data, and pedagogy.
Algorithmic Colonialism and the Appropriation of Indigenous Data
| Nouridin Melo | Preprints.org | 2025-05-19
Analyzes how algorithmic systems can appropriate, commodify, and misrepresent Indigenous knowledge in Cameroon's ethnolinguistic communities.
Digital Colonialism: How Social Media Enables New Violations of Cultural Rights
| Emese Ilyes | OpenGlobalRights | 2025-02-27
Explores how platforms commodify marginalized identities and shift control over cultural narratives away from the communities that sustain them.
An Intellectual History of Digital Colonialism
| Toussaint Nothias | Journal of Communication | 2025
Traces the historical precedents and Global South activist traditions that shaped contemporary theories of digital colonialism.
Digital Colonialism and the Role of Local Intermediaries
| J. O. Effoduh | Business and Human Rights Journal | 2025
Uses a TWAIL lens to examine how Big Tech and local intermediaries reinforce African data extraction, infrastructure dependence, and algorithmic bias.
Artificial Intelligence Bias and Digital Colonialism in Global South AI Governance
| Research Authors | ResearchGate | 2025
Examines how Global North data and values shape AI deployed in the Global South and proposes context-aware governance alternatives.
Moving Toward Truly Responsible AI Development in the Global AI Market
| Brookings Authors | Brookings Institution | 2024-10-24
Connects data annotation in developing countries to colonial extraction and calls for labor rights, transparency, and fair distribution of AI benefits.
Addressing Digital Colonialism: A Path to Equitable Data Governance
| Bitange Ndemo | UNESCO Inclusive Policy Lab | 2024-08-08
Calls for equitable data governance, stronger local institutions, and policies that prevent foreign platforms from extracting value without accountability.
Generative AI and Digital Neocolonialism in Global Education
| Matthew Nyaaba, Alyson Wright, and Gyu Lim Choi | arXiv | 2024-06-05
Explains how generative AI may impose Western curricula, cultural references, languages, and data-control regimes on non-Western education.
Digital Coloniality
| Jan Hendrik Kroeze | South African Journal of Information Systems | 2024
Defines digital coloniality through an Ubuntu-centered critique of information systems, technological dependence, and Western epistemic dominance.
Decolonizing LLMs: An Ethnographic Framework for AI in African Contexts
Offers an ethnographic framework for identifying and contesting digital colonialism in AI deployments across Ethiopia, Ghana, Kenya, Nigeria, and South Africa.
AI in the Global South: Opportunities and Challenges Toward More Inclusive Governance
| Landry Signé and colleagues | Brookings Institution | 2023-11-01
Surveys governance gaps that can expose Global South communities to imported bias, weak data protections, and unequal AI development.
AI Is Steeped in Big Tech's Digital Colonialism
| Khari Johnson | Wired | 2023-05-25
Profiles Abeba Birhane's work auditing web-scale datasets and challenging the Western corporate power embedded in AI development.
Risk and the Future of AI: Algorithmic Bias, Data Colonialism, and Marginalization
| Arora and colleagues | Information and Organization | 2023
Develops a relational-risk perspective for understanding how AI bias and data colonialism intensify the marginalization of already disadvantaged communities.
From Making Up Professionals to Epistemic Colonialism
| Dimitra Petrakaki and colleagues | Social Science and Medicine | 2023
Studies how digital health platforms transfer professional knowledge into Global South contexts and may impose external standards of expertise.
Digital Coloniality and Next Billion Users
| Tosin D. Oyedemi | Information, Communication and Society | 2021
Uses Google's Nigerian connectivity initiatives to reveal the market expansion and development narratives underlying digital coloniality.
Data Epistemologies, the Coloniality of Power, and Resistance
| Paola Ricaurte | Television & New Media | 2019-03-07
Explains how data systems reproduce the coloniality of power by imposing dominant ways of knowing while marginalizing alternative epistemologies.
Digital Colonialism: The 21st-Century Scramble for Africa
| Danielle Coleman | Michigan Journal of Race & Law | 2019
Examines how Western technology companies extract and control African user data while weak protections and infrastructure dependence limit local power.
Data Colonialism: Rethinking Big Data’s Relation to the Contemporary Subject
| Nick Couldry and Ulises A. Mejias | Television & New Media | 2018-09-02
Frames the continuous capture of human life through data as a colonial appropriation that enables discrimination, influence, and unlimited capitalization.
Wikipedia, Wikidata, and Knowledge Gaps
Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation
| Yannis Karmim and colleagues | arXiv | 2026-02-16
Uses Wikipedia and Wikidata to build a Latin American cultural benchmark showing uneven LLM knowledge across countries and languages.
Reducing AI Bias by Reducing Wiki Bias
| Wikimedia UK | Diff | 2026-02-07
Explains how Wikipedia's representation gaps are inherited by AI systems that rely on its openly licensed knowledge.
Waking Students Up to Systemic Bias: Using Wikipedia in Critical Pedagogy
| SAGE Authors | SAGE Journals | 2025-10-03
Describes how students can use Wikipedia editing to identify systemic omissions and participate in reshaping public knowledge.
Social Biases in Knowledge Representations of Wikidata Separate Global North from Global South
| Paramita Das, Sai Keerthana Karnam, Aditya Soni, and Animesh Mukherjee | arXiv | 2025-05-05
Shows that bias in Wikidata-based occupation predictions mirrors socioeconomic and cultural divisions between the Global North and Global South.
Wikipedia's Indian Problem: Settler Colonial Erasure of Native American Knowledge
| Kyle Keeler | Settler Colonial Studies | 2025
Details how Wikipedia's policies and editorial culture erase or distort Native American histories, philosophies, and knowledge systems.
Demographic Disparity in Wikipedia Coverage
| Y. Yu and colleagues | EPJ Data Science | 2025
Finds that biographies outside North America appear in fewer language editions and receive shorter coverage than North American subjects.
Whose Knowledge Is Valued? Epistemic Injustice in CSCW Applications
| Leah Hope Ajmani and colleagues | arXiv | 2024-07-03
Uses cases from online communities to show how platform rules can discount non-Western, experiential, and marginalized forms of knowledge.
Low-Resourced Languages and Online Knowledge Repositories
| Hellina Hailu Nigatu, John Canny, and Sarah E. Chasins | arXiv | 2024-05-26
Documents obstacles faced by Amharic, Afan Oromo, and Tigrinya contributors, including missing sources and poor language-technology support.
Assessing Knowledge Organization Systems from a Gender Perspective
| Miquel Centelles and colleagues | Journal of Documentation | 2024
Examines how classification systems and Wikipedia-related structures can reproduce gendered exclusions and barriers to representation.
Social Scientists Can't Ignore the Power of Wikipedia or Its Systemic Biases
| LSE Impact Blog Authors | LSE Impact Blog | 2023-04-06
Argues that researchers must engage with Wikipedia because its gaps and editorial patterns influence public understanding far beyond the platform.
Too Soon to Count? How Gender and Race Cloud Notability on Wikipedia
| Heather Ford and colleagues | Big Data and Society | 2023-03-29
Examines how notability disputes disproportionately challenge biographies of marginalized people and obscure racial inequality in Wikipedia research.
Wikipedia's Race and Ethnicity Gap and the Unverifiability of Whiteness
| Michael Mandiberg | Social Text | 2023-03-01
Shows how Wikipedia's race and ethnicity categorization debates reveal Eurocentric and anti-Black assumptions while treating whiteness as unmarked.
Diversity Matters: Robustness of Bias Measurements in Wikidata
| Paramita Das and colleagues | arXiv | 2023-02-27
Finds that bias measures vary across demographic groups and algorithms, challenging one-size-fits-all approaches to knowledge-graph fairness.
We Need a Woman in Music: Exploring Wikipedia's Values on Article Priority
| Mo Houtti and colleagues | arXiv | 2022-08-17
Studies how Wikipedia prioritization systems can either reinforce or reduce gender and geographic imbalance.
Wikipedia's Enlightenment Problem
| Matthew A. Vetter | Selected Papers of Internet Research | 2022
Argues that verifiability and notability rules privilege Enlightenment and colonial knowledge practices while excluding Indigenous sources.
Decolonizing Wikipedia
| Ian Ramjohn | Wikimedia Commons | 2022
Surveys Wikipedia's colonial knowledge structures and the movements working to diversify contributors, sources, and coverage.
Indigenous Knowledge on Wikipedia and Wikidata
| Ian Ramjohn | Wiki Education | 2021-11-23
Reflects on the difficulty of fitting relational Indigenous ways of knowing into Wikipedia prose and Wikidata's structured claims.
Celebrating Latinx Stories, Culture, and Contributions on Wikimedia Projects
| Jorge Vargas | Wikimedia Foundation | 2021-09-15
Discusses how colonial language, outsider narration, and contributor gaps shape Latin American representation across Wikimedia.
Analyzing Race and Country of Citizenship Bias in Wikidata
| Zaina Shaik, Filip Ilievski, and Fred Morstatter | arXiv | 2021-08-11
Finds overrepresentation of white Europeans and North Americans in Wikidata's STEM records and underrepresentation of other racial and national groups.
New Maps for an Inclusive Wikipedia
| Research Authors | ResearchGate | 2021
Proposes using decolonial scholarship and critical maps to counter the historical narratives and geographic imbalances embedded in Wikipedia.
Do Black Wikipedians Matter? Confronting the Whiteness in Wikipedia with Archives and Libraries
| Kai Alexis Smith | Wikipedia and Academic Libraries | 2021
Explores Black participation, racialized editing experiences, and how archives and libraries can challenge Wikipedia’s institutional whiteness.
The Positioning Matters: Estimating Geographical Bias in the Multilingual Record of Biographies on Wikipedia
| Pablo Beytía | Wiki Workshop | 2020
Measures how a person's birthplace affects the likelihood and extent of biographical coverage across Wikipedia language editions.
The Right Information: Perceptions of Information Bias Among Black Wikipedians
| Researchers | Journal of Documentation | 2019-09-03
Shows that racial identity and perceptions of information quality strongly influence why Black editors contribute to Wikipedia.
Wikipedia's World View Is Skewed by Rich, Western Voices
| James Temperton | Wired | 2015-09-15
Reports that editors from a small group of affluent Western countries dominate geographic knowledge production on Wikipedia.
Digital Divisions of Labour and Informational Magnetism
| Mark Graham, Ralph Straumann, and Bernie Hogan | Oxford Internet Institute | 2015-09-07
Maps how Wikipedia editing labor concentrates in wealthy countries and directs attention toward already well-represented places.
Wikipedia's Geography Problem: There Are More Articles About Antarctica Than Egypt
| Joseph Stromberg | Vox | 2014-09-14
Illustrates Wikipedia's geographic imbalance and explains why Europe and North America dominate the encyclopedia's mapped content.
Uneven Geographies of User-Generated Information
Shows that user-generated information can deepen informational poverty by concentrating content in already visible and wealthy locations.
Reliable Sources for Indigenous Knowledge: Dissecting Wikipedia's Catch-22
| Peter Gallert and Maja van der Velden | Wikimedia Commons | 2013
Explains how Wikipedia demands published sources while colonial publishing systems have historically excluded Indigenous knowledge.
Decentering Design: Wikipedia and Indigenous Knowledge
| Maja van der Velden | ResearchGate | 2013
Uses postcolonial computing to show how Wikipedia's templates, policies, and categories shape what counts as legitimate knowledge.
Wikipedia's Known Unknowns
| Mark Graham | The Guardian | 2009-12-02
Visualizes the severe underrepresentation of Africa and much of the Global South in Wikipedia's geotagged content.
Can History Be Open Source? Wikipedia and the Future of the Past
| Roy Rosenzweig | Journal of American History | 2006
Assesses Wikipedia as a historical knowledge system whose openness expands participation but also reflects contributors' interests and omissions.
Search Engines, Platforms, and Algorithmic Visibility
Empire Amplifier: Uncovering the Prioritization of Colonial Content on Platforms
| Nel Escher and colleagues | arXiv | 2026-04-30
Finds that YouTube recommendations steer Kyrgyz children toward Russian-language content even when they express a preference for Kyrgyz media.
Invisible Languages of the LLM Universe
| Research Authors | arXiv | 2025-10-13
Explains how missing digital infrastructure makes many widely spoken languages effectively invisible to AI-mediated information systems.
Colonial Biases and Systemic Issues in Automated Content Moderation Systems
| Farhana Shahid and colleagues | AAAI/ACM AIES | 2025
Shows how moderation tools fail low-resource languages because of colonial assumptions, cultural flattening, and weak institutional support.
Digital Colonialism Beyond Surveillance Capitalism
| Saskia Singler and colleagues | University of Essex Repository | 2024
Examines how Global North technical norms and development agendas shape Nigeria's emerging digital identification systems.
Technodiversity as the Key to Digital Decolonization
| Domenico Fiormonte | UNESCO Courier | 2023-03-31
Argues that decolonizing digital knowledge requires alternative technologies that preserve linguistic, cultural, and biological diversity.
The Workers Behind AI Rarely See Its Rewards
| Astha Rajvanshi | Time | 2023
Profiles a model that pays rural Indian language-data workers more fairly and gives them continuing value from the datasets they create.
Democracy Can Still End Big Tech's Dominance Over Our Lives
| Shoshana Zuboff | Time | 2022-05-05
Argues that democratic institutions must reclaim information infrastructure from surveillance-capitalist firms that control knowledge and communication.
Data Colonialism Brings About a New Social Order
| Nick Couldry and Ulises Mejias | IT for Change | 2022
Describes data colonialism as the continuous extraction of everyday life through platforms for profit and social control.
Building a Socialist Social Media Commons
| Dan Schiller | IT for Change | 2022
Links platform ownership to digital colonialism and proposes social ownership of communications infrastructure and knowledge networks.
Security Implications of Digitalization: The Dangers of Data Colonialism
| Matthias Stürmer, Jasmin Nussbaumer, and Pascal Stöckli | arXiv | 2021-07-04
Shows how governments and researchers become dependent on private platforms that control environmental and public-interest data.
Narratives and Counternarratives on Data Sharing in Africa
| Rediet Abebe and colleagues | arXiv | 2021-03-01
Challenges deficit narratives about African data and centers histories of extraction, unequal benefit, and locally grounded expertise.
Towards Buen Vivir with Data
| Paola Ricaurte and colleagues | IT for Change | 2021-01-25
Draws on Latin American traditions to propose liberation pedagogy, autonomous design, and ecological alternatives to extractive data systems.
Digital Cultural Colonialism: Measuring Bias in Google Arts and Culture
| Melissa Terras | Melissa Terras Blog | 2021-01-18
Reports that Google Arts and Culture's aggregated collections are not neutral and reproduce geographic and institutional bias.
A Modest Proposal to Save the World Through Tequiology
| Álvaro Ramírez | Rest of World | 2020-12-09
Highlights Indigenous and community-led digital practices that resist Western narratives, platform concentration, and linguistic exclusion.
The Digital Colonialism Behind .tv and .ly
| Emma Grey Ellis | Wired | 2020-02-07
Examines how wealthy foreign actors gained control over valuable country-code domains belonging to small or formerly colonized territories.
Google Has a Striking History of Bias Against Black Girls
| Safiya Umoja Noble | Time | 2018-03-26
Explains how commercial search ranking reproduced racist and sexualized stereotypes about Black girls and women.
Google, Democracy and the Truth About Internet Search
| Carole Cadwalladr | The Guardian | 2016-12-04
Investigates how search ranking, autocomplete, and platform manipulation can elevate racist misinformation and shape public reality.
Stereotypes in Search Engine Results
| Gabriel Magno, Camila Souza Araújo, Wagner Meira Jr., and Virgilio Almeida | arXiv | 2016-09-18
Shows that image search results reproduce language-linked beauty stereotypes that often differ from the populations of the countries searched.
The Book Stops Here
| Daniel H. Pink | Wired | 2005-03-01
Provides an early account of Wikipedia's open editing model and the unresolved tension between democratized participation, expertise, and reliability.
Can ICANN Make a Global Net?
| Kendra Mayfield | Wired | 2000-03-09
Documents early struggles over English dominance, multilingual domain names, and unequal participation in global internet governance.
Indigenous Data Sovereignty and Knowledge Rights
Understanding the CARE Principles
| Imperial College London Library | Open Access and Digital Scholarship Blog | 2026-02-09
Explains why Indigenous data governance must prioritize people, collective authority, responsibility, ethics, and community benefit.
A Framework for Kara-Kichwa Data Sovereignty in Latin America and the Caribbean
| WariNkwi K. Flores, KunTikzi Flores, Rosa M. Panama, and KayaKanti Alta | arXiv | 2026-01-10
Presents an Indigenous legal and relational framework for governing data as ancestral memory rather than a freely extractable commodity.
AI and Indigenous Data Sovereignty
| Issue Editors | Somatechnics | 2025-12-18
Introduces Indigenous-led approaches to AI that resist digital colonialism and protect collective rights over knowledge and data.
Digital Sovereignty or Digital Colonialism?
| Peace and Humanity | Peace and Humanity Monitor | 2025-09-11
Explains how Indigenous data sovereignty counters centuries of external collection, classification, and control of community knowledge.
Recognising Indigenous Data Sovereignty and Governance
| R. Marriott and colleagues | ScienceDirect | 2025
Discusses how CARE-based, relationship-centered governance can support Indigenous authority and equitable outcomes in data-intensive research.
Indigenous Scientists Are Fighting to Protect Their Data and Their Culture
| Justine Calma | The Verge | 2025
Reports on Indigenous researchers building alternative infrastructure and governance to protect cultural, environmental, and scientific data.
Free, Prior and Informed Consent in Digitalisation
| Association for Progressive Communications | APC | 2024-04-23
Argues that internet and technology governance must uphold Indigenous consent, land rights, environmental justice, and meaningful participation.
Tools to Support Indigenous Data Sovereignty and Cultural Authority
| Local Contexts | Zenodo | 2024
Describes digital labels and notices that let Indigenous communities communicate governance rules for cultural heritage and data.
New Report and Guidelines for Indigenous Data Sovereignty in AI Developments
| UNESCO | UNESCO | 2023-12-11
Calls for Indigenous self-determination, informed consent, privacy, intellectual property, and control over AI uses of cultural and linguistic data.
In Consideration of Indigenous Data Sovereignty: Data Mining as a Colonial Practice
| Jennafer Shae Roberts and Laura N. Montoya | arXiv | 2023-09-19
Applies CARE principles to show how data mining can repeat colonial extraction when Indigenous communities lack authority over technology development.
Decolonisation, Global Data Law, and Indigenous Data Sovereignty
| Jennafer Shae Roberts and Laura N. Montoya | arXiv | 2022-07-28
Proposes legal and economic protections against digital neocolonialism and links data governance to Indigenous environmental and cultural rights.
Operationalizing the CARE and FAIR Principles for Indigenous Data Futures
| Stephanie Russo Carroll and colleagues | Scientific Data | 2021
Shows how CARE and FAIR principles can be combined so that data stewardship produces tangible collective benefits and respects authority.
The CARE Principles for Indigenous Data Governance
| Stephanie Russo Carroll and colleagues | Data Science Journal | 2020
Defines Collective Benefit, Authority to Control, Responsibility, and Ethics as foundations for Indigenous-centered data governance.
Operationalising Indigenous Data Governance
| Stephanie Russo Carroll | Ada Lovelace Institute | 2020
Explains how institutions can move from abstract support for sovereignty to practical changes in data access, control, and benefit.
UX Design in Online Catalogs and Traditional Knowledge Labels
| Dana Reijerkerk | First Monday | 2020
Examines practical challenges in adding Indigenous Traditional Knowledge Labels to museum and library catalog interfaces.
The Passamaquoddy Reclaim Their Culture Through Digital Repatriation
| Anna Marks | The New Yorker | 2019-01-30
Describes a Native-first archival project that returns curatorial control over historic recordings to the Passamaquoddy community.
Preservation of Indigenous Culture Among Indigenous Migrants Through Social Media
| Khavee Agustus Botangen, Shahper Vodanovich, and Jian Yu | arXiv | 2018-02-27
Studies how Igorot diaspora communities use Facebook to exchange, practice, and revitalize cultural knowledge.
Tribal Archives, Traditional Knowledge, and Local Contexts
| Kimberly Christen | Journal of Western Archives | 2015
Explains how Indigenous archival management and Traditional Knowledge Labels can reshape control over digital cultural heritage.
Using Modern Technologies to Capture and Share Indigenous Astronomical Knowledge
| N. M. Nakata and colleagues | arXiv | 2014-09-04
Proposes community-controlled digital tools for preserving and sharing Indigenous astronomy under culturally appropriate access rules.
Maasai Music on iTunes? DRM as an Asset for Indigenous Groups
| Kim Zetter | Wired | 2009-07-27
Examines efforts to give Maasai communities tools to record, catalog, protect, and potentially license their songs, stories, and dances.
Language, Translation, and Online Knowledge Inequality
Global North-South Science Inequalities Due to Language and Funding Barriers
| R. Turba and colleagues | Peer Community Journal | 2026
Shows how English dominance and unequal funding reduce the visibility and policy influence of research from the Global South.
How Can AI Support Language Digitization and Revitalization?
| Stanford HAI Authors | Stanford Institute for Human-Centered AI | 2026
Reviews community-led approaches to building digital tools for under-resourced languages while protecting local priorities and ownership.
AI Diffusion in Low Resource Language Countries
| Amit Misra and colleagues | arXiv | 2025-11-04
Finds that poor AI performance in low-resource languages creates an independent barrier to equitable adoption.
Why African Languages and Knowledge Systems Matter in Online Governance
| CIPESA | CIPESA | 2025-10-24
Argues that digital governance and AI policy must include African languages and knowledge systems rather than treating English-language data as universal.
Studies Explore Challenges of AI for Low-Resource Languages
| Tech Brew Staff | Tech Brew | 2025-05-05
Summarizes research showing that generative AI progress remains concentrated in English and a small number of well-resourced languages.
Knowledge from Non-English-Language Studies Broadens Conservation Policy
| F. C. Serrano and colleagues | Journal of Applied Ecology | 2025
Shows that excluding non-English research narrows evidence bases and biases biodiversity policy and scientific datasets.
Winning the Language Divide with AI
| Kirti Vashee | Imminent | 2025
Reviews debates over whether AI can help preserve low-resource languages without repeating centralized, extractive development models.
Language Bias, Not Knowledge Shortfall, Underestimates South American Knowledge
| H. Bampi and colleagues | Global and Planetary Change | 2024
Argues that limited scientific recognition of South American knowledge is driven by publication-language bias rather than weak local expertise.
We Launched the State of the Internet's Languages Report
| Whose Knowledge? | Whose Knowledge? | 2022-03-31
Introduces a community-sourced report showing that only a small fraction of the world's languages have meaningful online presence.
Non-Dominant Languages in the Digital Landscape
Examines how dominant and colonial languages determine who can communicate, create content, and access services in digital spaces.
Summary Report: State of the Internet's Languages
Combines platform audits and community stories to document how language inequality structures access to digital knowledge.
Decolonizing Minority Language Technology
| Internet Languages Contributors | State of the Internet's Languages | 2020
Explores how minority-language technology can be designed around community needs instead of dominant-language assumptions.
Towards a Multilingual Internet
| Whose Knowledge? | Whose Knowledge? | 2019-11-11
Summarizes a gathering focused on linguistic justice, community technology, and the survival of marginalized languages online.
Cecilia Tuyuc and the Right of Indigenous Languages and Their Knowledge Online
| Claudia Pozo and Cecilia Tuyuc | Whose Voices? | 2019-10-15
Describes the challenges of building Kaqchikel and other Mayan-language knowledge on Wikipedia and social media.
Indigenous Languages: Knowledge and Hope
| Minnie Degawan | UNESCO | 2019-01-10
Explains how language loss also destroys ecological, cultural, and historical knowledge that dominant digital systems often overlook.
Measuring Linguistic Diversity on the Internet
| UNESCO Authors | UNESCO | 2005
Reviews the technical and conceptual difficulty of measuring which languages are represented online and how deeply they are supported.
Introduction: The Multilingual Internet
| Brenda Danet and Susan C. Herring | Journal of Computer-Mediated Communication | 2003
Introduces research on how diverse languages and scripts adapt to communication technologies originally shaped around English.
Digital Archives, Metadata, and Colonial Collections
Inclusive Collections, Inclusive Libraries
| Research Libraries UK | RLUK | 2026-06-10
Collects discussions on decolonizing how libraries select, describe, contextualize, and provide access to difficult or offensive collections.
Against the Illusion: The Limits of Digital Repatriation in Restitution Debates
| Center for Art Law | Center for Art Law | 2025-12-08
Warns that digital copies can become substitutes for returning physical objects and may leave colonial ownership structures untouched.
Colonial Archives and Meaningful Digital Infrastructure
| Colonial Collections Consortium | Colonial Collections | 2025-01-24
Explores how digital access, enrichment, and infrastructure can expose rather than conceal the biases of colonial records.
Building Better Archival Futures by Recognizing Epistemic Injustice
| Charles Jeurgens | Boletim do Arquivo da Universidade de Coimbra | 2025
Uses epistemic injustice to identify harms embedded in colonial archives and guide more accountable digital archival practice.
Beyond Access: Rethinking Ownership, Justice, and Decolonization in Digital Repatriation
| B. A. Mohammed | Heritage Management Organization | 2025
Argues that decolonial digital repatriation must center the sovereignty, authority, and cultural values of source communities.
Metadata and Linked Open Data in Digital Heritage for Decolonization
| Reference Work Authors | Springer | 2024-12-28
Reviews how digital heritage projects can expose colonial practices and support more plural historical narratives.
Digitization Is Not Decolonization
| L. Gibson | Museum Worlds | 2024
Argues that putting colonial collections online does not by itself change ownership, interpretation, institutional power, or community authority.
Digital History and the Politics of Digitization
| Gerben Zaagsma | Digital Scholarship in the Humanities | 2023
Provides a framework for analyzing how funding, institutions, selection, and infrastructure shape what historical materials become digitally visible.
What's the Use of the Archive? Locality, Accessibility, and Digitalisation
| Larissa Schulte Nordholt and contributors | Boasblogs | 2022-04-05
Examines unequal archival access between the Global North and Global South and the limits of digitization as a remedy.
Digital Repatriation as a Decolonizing Practice in the Archaeological Archive
| Krystiana L. Krupa and Kelsey T. Grimm | Across the Disciplines | 2021
Argues that returning digital archival materials can support decolonization when descendant communities govern access, context, and reuse.
What's in a Name? Cataloguing in a Decolonising Library
| Institute of Historical Research Staff | Institute of Historical Research | 2020-05-15
Describes efforts to replace outdated and offensive subject terms so library catalogs better represent Indigenous peoples and histories.
The Crying Child: On Colonial Archives, Digitization, and Ethics of Care
| Temi Odumosu | Current Anthropology | 2020
Uses a colonial photograph to show the unresolved ethical harms of digitizing and recirculating images of colonized and enslaved people.
Paradoxes of Curating Colonial Memory
| Charles Jeurgens and Michael Karabinos | Archival Science | 2020
Shows how digital remixing can challenge colonial archives while still relying on records created through colonial power.
Building Critical Decolonial Digital Archives
| Bibhushana Poudyal | Xchanges | 2018
Outlines methods for building digital archives that acknowledge complexity, positionality, and the colonial politics of preservation.
Decolonising Archives
| L'Internationale Online | L'Internationale Online | 2016
Collects essays on recovering the political potential of archives and challenging institutional narratives inherited from colonialism.
Decolonizing the Internet and Knowledge Justice
VisibleWikiWomen 2026: Whose Data Tells Our Stories?
| Whose Knowledge? | Whose Knowledge? | 2026
Connects visual representation on Wikimedia to broader struggles over data, memory, consent, and who is authorized to tell public stories.
VisibleWikiWomen 2025
| Whose Knowledge? | Whose Knowledge? | 2025
Calls for freely licensed images of Black, Brown, Indigenous, trans, and nonbinary people to correct online visibility gaps.
Decolonizing Structured Data: A New Season of Whose Voices?
| Whose Knowledge? | Whose Knowledge? | 2024-10-21
Introduces conversations about how databases and structured knowledge systems encode hierarchy, absence, and institutional power.
Deep Dive into Decolonizing Structured Data
| Whose Knowledge? | Whose Knowledge? | 2023
Collects reflections on how structured data can reproduce colonial categories and how communities might redesign it around plural knowledge systems.
Can You Get Reliable Information About Safe Abortion in Your Language Online?
| Whose Knowledge? | Whose Knowledge? | 2022-12-09
Uses reproductive-health information to show how language inequality determines whether people can find trustworthy and life-saving knowledge online.
Challenging Epistemic Injustice Through Feminist Practice
| Camille E. Acey and colleagues | Whose Knowledge? | 2022-11-25
Examines how feminist organizing can confront the exclusion of marginalized people from public online knowledge.
Arya Jeipea on Our Existence Is Our Truth
| Arya Jeipea Karijo and Whose Knowledge? | Whose Knowledge? | 2022-06-23
Explores Kenyan transgender representation and the difficulty of making marginalized lived experience legible within dominant online knowledge systems.
Decolonizing Knowledge, Decolonizing the Internet
| Whose Knowledge? | Whose Knowledge? | 2022
Explains why internet reform must address power over knowledge production, not merely expand technical access.
Decolonizing the Internet: A Conversation with Whose Knowledge?
| Whose Knowledge? | Whose Knowledge? | 2022
Discusses strategies for centering marginalized communities in online history, representation, language, and infrastructure.
Whose Knowledge Is Online? Practices of Epistemic Justice for a Digital New Deal
| Az Causevic and Anasuya Sengupta | Whose Knowledge? | 2020-10-30
Defines online epistemic injustice and proposes community leadership, contextual design, and translocal solidarity as decolonial practices.
Seeing Is Believing: Why Online Visibility Matters
| Mariana Fossatti and Adele Vrana | Whose Knowledge? | 2020-03-09
Connects the visual erasure of African women to broader gaps in Wikipedia and the public internet.
Beyond Internet Access: Seeking Knowledge Justice Online
| Whose Knowledge? | Whose Knowledge? | 2020
Argues that human-rights approaches must address whose knowledge is visible and authoritative, not only who can connect to the internet.
New Resource: Adding Our Knowledge to Wikipedia
| Whose Knowledge? | Whose Knowledge? | 2018-11-22
Explains how Native American communities approached Wikipedia while navigating consent, reliability rules, and community authority.
Transformative Practices for Sharing Marginalized Knowledge
| Whose Knowledge? | Whose Knowledge? | 2018-11-19
Offers community-based practices for documenting and sharing knowledge without reproducing extractive relationships.
Snapshot of Our Decolonize the Internet Conference
| Whose Knowledge? | Whose Knowledge? | 2018-08-27
Summarizes a Global South-centered gathering on whose languages, histories, and infrastructures shape the internet.