Research, as evidence rather than reading

Papers and institutional reports recorded as inputs to this record, each labelled with its venue. A preprint is a preprint: not peer reviewed, and not treated as established fact because it is published.

31 papers and reports, of which 22 are preprints.
Explore this data →
31papers and reports recorded
22preprints, not peer reviewed
0peer-reviewed articles
9institutional reports

Research enters this record as evidence, not as decoration: the exposure measures on the Jobs page and the affordability layer on the Access page are built on papers that appear below. Each item states its venue, because a preprint and a peer-reviewed article are not the same claim.

Preprints

Recorded from arXiv and similar servers. A preprint has not been peer reviewed, and nothing here is treated as established fact because it appears on a preprint server. Where a paper underpins an indicator, that is stated with the indicator.

  • PREPRINTPolicy and regulationInfrastructure

    Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech

    arXiv

    Andrew Anogie Uduimoh; Hadiza Umar Yusuf; Oluwafemi Osho · 2026-09-21

    Country named: NGA

    Abstract, in the authors’ words

    Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulatory instrument specifies what pre-deployment evaluation such systems must undergo before procurement. This paper reviews African fintech AI governance across global, continental, and Nigerian instruments, and shows that safety is affirmed as a principle while pre-deployment evaluation is operationally unspecified. Generic safety benchmarks cannot surface the failure modes most relevant to this domain, since none contain Nigerian institutional content or test for false positive misclassification of legiti…

  • PREPRINTPolicy and regulationInfrastructure

    Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability

    arXiv

    Sharifa Sultana; Syed Ishtiaque Ahmed · 2026-09-20

    Abstract, in the authors’ words

    Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for a…

  • PREPRINTSafety and governance of AIJobs

    When AI Enters the Workplace, Who Faces Greater Risks? A Gendered Analysis

    arXiv

    Miriam Fernandez; Ángel Pavón Pérez; Damiano Giallongo; Davide Ghia; Maryam Yaqub; Daniele Quercia · 2026-09-18

    Abstract, in the authors’ words

    Gender inequality remains a persistent structural feature of the labour market, shaping women's lifetime earnings and economic security. As artificial intelligence (AI) transforms organisational practices, there is growing concern that existing disparities may be unintentionally amplified through task automation, unequal access to upskilling opportunities, and differential returns obtained from technological change. In this paper, we examine how exposure to AI-driven innovation varies across male- and female-dominated occupations, with particular attention to differences across the skill and wage distribution. Using a novel dataset that links occupational characteristics to measures of AI ex…

  • PREPRINTEmployment and outsourcingJobs

    Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria

    arXiv

    Abbas M. Rabiu; Abdulrazaq A. Zubair; Um-mulkhairi Ibrahim; Tolulope Olusuyi; Shaheeda Farouq; Safwan M. Dafi · 2026-09-16

    Country named: NGA

    Abstract, in the authors’ words

    Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness, particularly in low- and middle-income countries (LMICs) where infrastructural and training gaps persist. This cross-sectional study evaluated awareness, attitudes, preparedness, and barriers to AI adoption among 761 healthcare professionals across multiple disciplines and practice settings in Nigeria. Data were collected between December 2025 and March 2026 using a structured, validated questionnaire. Overall awareness of AI in healthcare was high (92.6%); however, objective knowledge and self-reported preparedness remaine…

  • PREPRINTData centres and computeInfrastructure

    ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters

    arXiv

    Rui Lu; Rui Ge; Huanghuang Liang; Xiaobo Zhou; Dan Wang · 2026-09-14

    Abstract, in the authors’ words

    Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energy-efficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts per-GPU frequency and micro-batch size during executi…

  • PREPRINTDevices and hardwareAccess

    BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices

    arXiv

    Maoliang Li; Hailong Zou; Taohong Han; Haoze Chi; Jiayu Chen; Zihao Zheng · 2026-09-13

    Abstract, in the authors’ words

    Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse computation stalled on weight transfers. Each transfer serves few tokens before execution moves on. We exploit the multi-token verification window of speculative decoding to decouple expert movement from single-token execution, enabling weight reuse, contiguous flash reads, and load-compute overlap. We present \textsc{BigMoMo}, a mobile MoE runtime that exploits this window across the memory hierarchy.…

  • PREPRINTModels and capabilityInfrastructure

    HeRo: History-Aware Routing for Efficient LLM Inference

    arXiv

    Hongjin Lin; Wentao Wan; Keze Wang · 2026-09-08

    Abstract, in the authors’ words

    Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, however, treat each routing decision as a local operation conditioned solely on the current hidden state which is a formulation that overlooks the sequential, path-dependent nature of routing across depth: earlier decisions shape the representations seen by downstream routers, and the layer-usage objective couples all decisions jointly. We propose History-Aware Routing (HeRo), a dynamic routing framework that resolves this mismatch by introducing a router memory mechanism to maintain an explicit routing state across model depth. The memory is co…

  • PREPRINTModels and capabilityInfrastructure

    A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

    arXiv

    Maysam Khatib; Moysis Symeonides; Demetris Trihinas; George Pallis; Marios D. Dikaiakos · 2026-09-08

    Abstract, in the authors’ words

    Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes. We evaluate multiple open-weight LLMs and quantization variants using a fixed question-answering workload, and compare them against GPT-4o as a cloud-hosted accuracy and latency reference. Our benchmarking pipeline reports accuracy, model footprint, per-token decoding latency, …

  • PREPRINTPolicy and regulationInfrastructure

    Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management

    arXiv

    Jiaxun Lu; Xiang Zhang; Yunfeng Shao · 2026-09-07

    Abstract, in the authors’ words

    Large language model (LLM) inference increasingly spans models that differ in size, capability, price, and provider. This shift creates two costs for developers. One is the integration cost of choosing among and switching between many models. The other is the inference cost of rebuilding a KV cache when it is unavailable or incompatible with the selected model. We define and analyze the Unified AI Gateway as a system setting for an edge-deployed AI traffic hub. It coordinates model routing, KV cache management, and compute placement across end devices, edge resources, and cloud model services. At request time, the gateway jointly selects a target model, an execution site, and a KV cache acti…

  • PREPRINTDevices and hardwareInfrastructure

    Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening

    arXiv

    Rui Xiao; Yili Xu · 2026-08-31

    Abstract, in the authors’ words

    Recent advances in large language models (LLMs) have demonstrated exceptional performance in protein-ligand interaction prediction, but state-of-the-art pipelines for large-scale virtual screening almost exclusively rely on high-end GPU clusters with hundreds of gigabytes of memory, creating prohibitive hardware barriers for small academic teams. In this work, we present a fully local low-resource framework that deploys the 175-billion-parameter DeepSeek 175B LLM on a single consumer-grade RTX 4060 laptop equipped with 32GB system RAM and 8GB VRAM, completing a full 200k-scale protein-ligand virtual screening workflow across 20 distinct protein targets. Our implementation achieves 100x throu…

  • PREPRINTSkills and talentJobs

    From Producing to Validating: How AI Is Deskilling Freelancers

    arXiv

    Nakul Rajpal · 2026-08-26

    Abstract, in the authors’ words

    Generative AI is promoted as a way to enhance knowledge work, yet its benefits and drawbacks fall unevenly across the workforce. Freelance and gig workers, who commonly lack the upskilling pathways available to traditional employees, face heightened risks to both skill development and job security as AI adoption advances. We review empirical evidence on AI's impact on knowledge-worker workflows and upskilling, then predict the primary and downstream effects of AI adoption among clients and workers in the freelance economy. We anchor this in two cases of the same shift, machine-translation post-editing and software development. We argue that freelancers are the leading edge of a change that a…

  • PREPRINTAccess and affordabilityAccess

    TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

    arXiv

    Milan Gritta; Patrik Lambert; Jihye Back; Amril Nazir · 2026-08-19

    Abstract, in the authors’ words

    The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African language machine translation, while the lack of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models (SLMs). We introduce TranslatePsy-AfriSLM, a collection of open-source machine translation resources for 19 Sub-Saharan African languages, including curated parallel data, African-specialized synthetic data, and a family of fine-tuned SLMs. Our empirical study shows that unified quality-estimation filtering remo…

  • PREPRINTModels and capabilityJobs

    CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

    arXiv

    Pattaraphon Kenny Wongchamcharoen; Kris Gulati; Min Min Fong; Abhishek Nagaraj · 2026-08-19

    Abstract, in the authors’ words

    Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are …

  • PREPRINTSkills and talentJobs

    Stranded credentials: how a skill-signaling market absorbed generative AI

    arXiv

    Song Yao · 2026-08-17

    Abstract, in the authors’ words

    Platforms summarize providers' past achievements into credentials that buyers use to judge quality. Generative AI can now produce much of the work those achievements certify, raising fears that those quality signals are worthless. We audit how well such credentials stay informative in Kaggle's 2010-2026 archive, where medals are won on predictions scored against withheld answers and two evaluation formats ran side by side. Across 444,698 participations, a medal's power to predict performance sits almost entirely in its first year, in both formats and eras. Fresh medals kept most of their value through the AI transition. About half of the collapse in the informativeness of one format's medals…

  • PREPRINTPricing and affordabilityInfrastructure

    The Embedder's Dilemma: LLMs Are Better, but at What Cost?

    arXiv

    Adnan El Assadi; Niklas Muennighoff; Jinhyuk Lee · 2026-08-13

    Abstract, in the authors’ words

    Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM c…

  • PREPRINTModels and capabilityInfrastructure

    Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

    arXiv

    Zixuan Lan; Yanhong Li; Jiawei Zhou · 2026-08-13

    Abstract, in the authors’ words

    Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights. Under a simple retention-ratio control, RMM provides a smooth and predictable accuracy-efficiency trade-off. Across language models ranging from 1B to 70B parameters, we find that reduction tolerance depends on the model family, task, component, and retention ratio, although it often improves with model s…

  • PREPRINTPricing and affordabilityInfrastructure

    Every Time I Hire a Linguist, Inference Costs Go Down: Linguistic Rules as Effective Prompt Compressors

    arXiv

    Jianfei Ma; Zhaoxin Feng; Emmanuele Chersoni; Si Chen · 2026-07-28

    Abstract, in the authors’ words

    Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can \textbf{linguistic rules alone} serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting lingui…

  • PREPRINTModels and capabilityInfrastructure

    Inference Economics of Enterprise Coding Agents: Cloud vs. On-Premise LLMs

    arXiv

    Sheng-Wei Peng; Yi-Hsun Lin; Yi-Pei Lee · 2026-07-13

    Abstract, in the authors’ words

    Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost -- and on-premise quantized open-weights models, which promise low-marginal-cost scaling and data sovereignty at some loss of reasoning fidelity. We study this trade-off through a single-developer, non-randomized longitudinal case study over two contiguous 28-day periods on a production monorepo: an API-based Claude Opus 4.7/4.8 configuration using Claude Code versus an on-premise GLM-5.1/5.2 configuration using Opencode, quantized to NVFP4, on NVIDIA Blackwell hardware. Analyzing LLM telemetry and Git history, we find that prompt caching (99.3% hit rate…

  • PREPRINTModels and capabilityAccess

    Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 — A quantitative scenario analysis of inference economics

    arXiv:2607.07207

    Satoshi Matsuoka · 2026-07-08

    Abstract, in the authors’ words

    We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local runtimes), and the entry of Meta and xAI into compute resale on fleets bought before the memory repricing. Formulating inference economics in dollars per petabyte of bandwidth delivered (\$/PB) -- model-agnostic for bandwidth-bound decode -- we show the entrant-incumbent cost gap never closes: a depreciation conveyor delivers newly amortized fleets to incumbents faster than hardware prices normalize (3.2x in 2026, 1.9x in 2027, re-widening to 3-4x by 2029-30)…

  • PREPRINTPolicy and regulationInfrastructure

    Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

    arXiv (Hung et al.)

    Kai-Hsin Hung; Sumaya Nur Adan; Krupa Suchak; Armita Sadeghian Barzoki; Kofi Yeboah; Mohammad Amir Anwar · 2026-06-24

    Abstract, in the authors’ words

    Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute primarily as a technical input rather than as an outcome of investment, ownership, and financial control. This paper examines AI infrastructure investment flows across Africa through a systematic analysis of 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, we analyze who invests in AI-relevant infrastructure and where investments concentrate. Our findings reveal a highly concentrated landscape dominated by global data center operators, hyperscale technology firms, and development fina…

  • PREPRINTDevices and hardwareAccess

    LLM Inference at the Edge: Mobile, NPU and GPU Performance

    arXiv:2603.23640

    Pranay Tummalapalli; Sahil Arayakandy; Ritam Pal; Kautuk Kundan · 2026-03-24

    Abstract, in the authors’ words

    Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power, thermal envelope, and memory. We benchmark Qwen 2.5 1.5B (4-bit quantised) across four platforms: a Raspberry Pi 5 with Hailo-10H NPU, a Samsung Galaxy S24 Ultra, an iPhone 16 Pro, and a laptop NVIDIA RTX 4050 GPU. Using a fixed 258-token prompt over 20 warm-condition iterations per device, we measure throughput, latency, power, and thermal behaviour. For mobile platforms, thermal management supersedes peak compute as the primary constraint: the iPhone 16 Pro loses nearly half its throughput within two iterations, and the S24 Ultra suffers a hard OS-…

  • PREPRINTModels and capabilityAccess

    Beyond Benchmarks: The Economics of AI Inference

    arXiv:2510.26136

    Boqin Zhuang; Jiacheng Qiao; Mingqian Liu; Mingxing Yu; Ping Hong; Rui Li · 2025-10-30

    Abstract, in the authors’ words

    The inference cost of Large Language Models (LLMs) has become a critical factor in determining their commercial viability and widespread adoption. This paper introduces a quantitative ``economics of inference'' framework, treating the LLM inference process as a compute-driven intelligent production activity. We analyze its marginal cost, economies of scale, and quality of output under various performance configurations. Based on empirical data from WiNEval-3.0, we construct the first ``LLM Inference Production Frontier,'' revealing three principles: diminishing marginal cost, diminishing returns to scale, and an optimal cost-effectiveness zone. This paper not only provides an economic basis …

Institutional research

Reports published by an institution rather than a journal: statistical agencies, development banks, standards bodies. Not peer reviewed, and not a preprint either.

How items get here

Classified by source, checked by rule

A signal is recorded here when its source is a paper or an institution. That classification is mechanical and the rules are published in the open: a URL from a preprint server or a journal, or from an institution we track, determines the venue. Topic and country are matched from the headline and can be wrong; they narrow a search, they are not a claim.

What this layer does not do

It does not peer review, rank or endorse. It does not paraphrase findings: abstract text is quoted and attributed. Where a paper matters enough to explain, the explanation is written, reviewed, and then published as InferenceAfrica’s reading, never presented as the authors’ conclusion.

Explore this data

Compare records, build a chart and download your selection in the Data Explorer.

Open Data Explorer →