The Earth seen from space, a thin lit crescent against the dark

Welcome to the Internet

A lecture series on how the infrastructure of the internet is changing — and why that matters to anyone who reads the web, publishes on it, takes data from it, or builds agents that go out into it. Coming soon.

We talk a great deal about what AI does. We talk much less about the environment it evolves alongside. Models are trained on the web, search it, cite it, and increasingly stand between it and the reader — and the web is changing in response, in what gets crawled, what stays indexed, which interfaces remain open, and what is therefore reachable at all. That second half of the pair deserves the same attention as the first, because it is where the assumptions sit.

The internet is a large shadow on the wall of Plato’s cave. It is a model of the world, and a strange one: everything that reaches you through it has passed through decisions about what was published, what was collected, what was ranked, and what was shown. Those decisions are the assumptions of the model. Reading anything off the web with the right amount of trust means knowing what they are, and noticing when they change.

I wrote my master’s thesis on these systems. That does not make me right about every detail, or my experiment methodologically flawless, much as I worked at it. It does mean I have read most of the current research literature on the subject, and these lectures are an attempt to hand on what is in it.

The lectures will be published here in English, free and open to anyone.

01

Bibliography

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. KDD 2024. arXiv:2311.09735.
  2. Aumüller, M., Bernhardsson, E., & Faithfull, A. (2020). ANN-Benchmarks: A Benchmarking Tool for Approximate Nearest Neighbor Algorithms. Information Systems, 87.
  3. Baeza-Yates, R., & Ribeiro-Neto, B. (2011). Modern Information Retrieval (2nd ed.).
  4. Bajaj, P. et al. (2016). MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268.
  5. Bruch, S. (2024). Foundations of Vector Retrieval. Springer. arXiv:2401.09350. DOI: 10.1007/978-3-031-55182-6.
  6. Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1), 37–46. DOI: 10.1177/001316446002000104.
  7. Cotton, I. W., & Greatorex, F. S., Jr. (1968). Data Structures and Techniques for Remote Computer Graphics. AFIPS Fall Joint Computer Conference (San Francisco, 9–11 December 1968). DOI: 10.1145/1476589.1476661. First recorded use of the term API.
  8. Craswell, N., Hawking, D., Wilkinson, R., & Wu, M. (2003). Overview of the TREC 2003 Web Track. TREC 2003 (NIST SP 500-255).
  9. Craswell, N. et al. (annual, since 2019). Overview of the TREC Deep Learning Track.
  10. Croft, W. B., Metzler, D., & Strohman, T. (2010). Search Engines: Information Retrieval in Practice.
  11. Cronen-Townsend, S., & Croft, W. B. (2002). Quantifying Query Ambiguity. HLT 2002. DOI: 10.3115/1289189.1289266.
  12. Dai, S. et al. (2024). Neural Retrievers are Biased Towards LLM-Generated Content. KDD 2024. arXiv:2310.20501.
  13. Efron, B. (1987). Better Bootstrap Confidence Intervals. JASA, 82(397), 171–185. DOI: 10.1080/01621459.1987.10478410.
  14. Ernst, M. D. (2004). Permutation Methods: A Basis for Exact Inference. Statistical Science, 19(4), 676–685. DOI: 10.1214/088342304000000396.
  15. Field, C. A., & Welsh, A. H. (2007). Bootstrapping Clustered Data. JRSS: Series B, 69(3), 369–390. DOI: 10.1111/j.1467-9868.2007.00593.x.
  16. Formal, T., Piwowarski, B., & Clinchant, S. (2021). SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. SIGIR 2021. arXiv:2107.05720.
  17. Friedman, M. (1937). The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance. JASA, 32(200), 675–701. DOI: 10.1080/01621459.1937.10503522.
  18. Furnas, G. W., Landauer, T. K., Gomez, L. M., & Dumais, S. T. (1987). The Vocabulary Problem in Human-System Communication. Communications of the ACM, 30(11), 964–971. DOI: 10.1145/32206.32212.
  19. Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997.
  20. Comanici, G. et al. (Google DeepMind Gemini Team) (2025). Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv:2507.06261. 2,600+ co-authors.
  21. Guo, J. et al. (2020). A Deep Look into Neural Ranking Models for Information Retrieval. Information Processing & Management, 57(6).
  22. Guo, J. et al. (2022). Semantic Models for the First-stage Retrieval: A Comprehensive Review. ACM TOIS, 40(4).
  23. Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
  24. Huang, P.-S. et al. (2013). Learning Deep Structured Semantic Models for Web Search using Clickthrough Data. CIKM 2013.
  25. Huang, J.-T. et al. (2020). Embedding-based Retrieval in Facebook Search. KDD 2020. arXiv:2006.11632.
  26. Indyk, P., & Motwani, R. (1998). Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality. STOC 1998.
  27. Izacard, G. et al. (2022). Unsupervised Dense Information Retrieval with Contrastive Learning. TMLR. arXiv:2112.09118.
  28. Järvelin, K., & Kekäläinen, J. (2002). Cumulated Gain-based Evaluation of IR Techniques. ACM TOIS, 20(4), 422–446. DOI: 10.1145/582415.582418.
  29. Jégou, H., Douze, M., & Schmid, C. (2011). Product Quantization for Nearest Neighbor Search. IEEE TPAMI, 33(1).
  30. Johnson, J., Douze, M., & Jégou, H. Billion-scale Similarity Search with GPUs (FAISS). IEEE Transactions on Big Data. arXiv:1702.08734.
  31. Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020, 6769–6781. arXiv:2004.04906.
  32. Khattab, O., & Zaharia, M. (2020). ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. SIGIR 2020. arXiv:2004.12832.
  33. Lee, K., Chang, M.-W., & Toutanova, K. (2019). Latent Retrieval for Weakly Supervised Open Domain Question Answering (ORQA). ACL 2019. arXiv:1906.00300.
  34. Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. arXiv:2005.11401.
  35. Lin, J., Nogueira, R., & Yates, A. (2021). Pretrained Transformers for Text Ranking: BERT and Beyond. Synthesis Lectures on Human Language Technologies, Morgan & Claypool. arXiv:2010.06467. DOI: 10.2200/S01123ED1V01Y202108HLT053.
  36. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of the ACL: EMNLP 2023, 7001–7025. arXiv:2304.09848.
  37. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. TACL.
  38. Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026). arXiv:2607.14035.
  39. Ma, X., Gong, Y., He, P., Zhao, H., & Duan, N. (2023). Query Rewriting for Retrieval-Augmented Large Language Models. EMNLP 2023. arXiv:2305.14283.
  40. MacAvaney, S. et al. (2022). ABNIRML: Analyzing the Behavior of Neural IR Models. TACL.
  41. Malkov, Yu. A., & Yashunin, D. A. (2018/2020). Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs (HNSW). IEEE TPAMI, 42(4), 824–836. arXiv:1603.09320. DOI: 10.1109/TPAMI.2018.2889473.
  42. Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
  43. Mitra, B., & Craswell, N. (2018). An Introduction to Neural Information Retrieval. Foundations and Trends in Information Retrieval, 13(1), 1–126. DOI: 10.1561/1500000061.
  44. Muennighoff, N. et al. (2023). MTEB: Massive Text Embedding Benchmark. EACL 2023. arXiv:2210.07316.
  45. Nigam, P. et al. (2019). Semantic Product Search. KDD 2019.
  46. Nogueira, R., & Cho, K. (2019). Passage Re-ranking with BERT. arXiv:1901.04085.
  47. Ogilvie, P., & Callan, J. (2003). Combining Document Representations for Known-Item Search. SIGIR '03, 143–150. DOI: 10.1145/860435.860463.
  48. Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does Conversational SEO Work? arXiv:2506.11097.
  49. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP-IJCNLP 2019, 3982–3992. arXiv:1908.10084.
  50. Reimers, N., & Gurevych, I. (2021). The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes. ACL-IJCNLP 2021.
  51. Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. DOI: 10.1561/1500000019.
  52. Santhanam, K. et al. (2022). ColBERTv2. NAACL 2022. arXiv:2112.01488.
  53. Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't Measure Once: Measuring Visibility in AI Search (GEO). arXiv:2604.07585.
  54. Sciavolino, C., Zhong, Z., Lee, J., & Chen, D. (2021). Simple Entity-Centric Questions Challenge Dense Retrievers (EntityQuestions). EMNLP 2021, 6138–6148. arXiv:2109.08535.
  55. Sparck Jones, K. (1972). A Statistical Interpretation of Term Specificity and Its Application in Retrieval. Journal of Documentation, 28(1), 11–21. DOI: 10.1108/eb026526.
  56. Su, H., Yen, H., Xia, M., Shi, W., Muennighoff, N., et al. (2024). BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval. ICLR 2025. arXiv:2407.12883.
  57. Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., & Gurevych, I. (2021). BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets & Benchmarks 2021. arXiv:2104.08663.
  58. Voorhees, E. M. (1999). The TREC-8 Question Answering Track Report. TREC-8, NIST SP 500-246.
  59. Wang, W., Guo, Y., Shen, C., Ding, S., Liao, G., Fu, H., & Prabhakar, P. K. (2023). Integrity and Junkiness Failure Handling for Embedding-based Retrieval: A Case Study in Social Network Search. SIGIR '23. arXiv:2304.09287. DOI: 10.1145/3539618.3591831.
  60. Weller, O., Boratko, M., Naim, I., & Lee, J. (2025). On the Theoretical Limitations of Embedding-Based Retrieval (LIMIT). arXiv:2508.21038.
  61. Wilcoxon, F. (1945). Individual Comparisons by Ranking Methods. Biometrics Bulletin, 1(6), 80–83. DOI: 10.2307/3001968.
  62. Xiong, L. et al. (2021). Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval (ANCE). ICLR 2021. arXiv:2007.00808.
  63. Zhang, K., He, X., & Yao, J. (2026). From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms. arXiv:2604.25707.
  64. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629.
  65. Zhao, W. X., Liu, J., Ren, R., & Wen, J.-R. (2022). Dense Text Retrieval based on Pretrained Language Models: A Survey. arXiv:2211.14876.
  66. Zhu, Y. et al. (2023). Large Language Models for Information Retrieval: A Survey. arXiv:2308.07107.