Files
papers-we-love_papers-we-love/sublinear_algorithms/README.md
Toyeshh Medikonda 9b92363eac Fix broken hosted-paper link and typos in topic READMEs (#894)
* Fix broken hosted-paper link and typos in topic READMEs

The FRP index links the hosted paper as
[Deprecating the Observer Pattern](deprecating-the observer-pattern.pdf).
The file on disk really is named with a space, but an unescaped space is not
allowed in a CommonMark link destination, so the entry renders as literal text
instead of a link. Percent-encode the space so it resolves to the committed PDF.

Also correct misspellings in paper titles and descriptions:

- Algorithims -> Algorithms (comp_sci_fundamentals_and_history)
- simplifed -> simplified (data_structures)
- sucessful -> successful, numberator -> numerator,
  demoninator -> denominator, authoratitative -> authoritative
  (information_retrieval)
- Expresions -> Expressions (logic_and_programming)
- represention -> representation (mathematics)
- Probablistic -> Probabilistic (robotics, sublinear_algorithms)
- Philppe -> Philippe (streaming_algorithms)

* Rename the Deprecating the Observer Pattern PDF to drop the space

Per review, renaming the file rather than percent-encoding the link.
2026-09-29 05:45:39 -04:00

17 lines
1.8 KiB
Markdown

# Sublinear Algorithms
## Hosted Papers :open_file_folder:
* :scroll: **[Probabilistic Counting Algorithms for Database Applications](https://github.com/papers-we-love/papers-we-love/blob/main/sublinear_algorithms/1985-Flajolet-Probabilistic-counting.pdf)**
This paper introduces a class of probabilistic counting algorithms with which one can estimate the number of distinct elements in a large collection of data (typically a large file stored on disk) in a single pass using only a small additional storage (typically less than a hundred binary words) and only a few operations per element scanned. The algorithms are based on statistical observations made on bits of hashed values of records. They are by construction totally insensitive to the replicative structure of elements in the file; they can be used in the context of distributed systems without any degradation of performances and prove especially useful in the context of data bases query optimisation
*Flajolet, Philippe, and G. Nigel Martin. "Probabilistic counting algorithms for data base applications." Journal of computer and system sciences 31.2 (1985): 182-209.*
* :scroll: **[An Elementary Proof of a Theorem of Johnson and Lindenstrauss](https://github.com/papers-we-love/papers-we-love/blob/main/sublinear_algorithms/An-Elementary-Proof-of-a-Theorem-of-Johnson-and-Lindenstrauss.pdf)**
A result of Johnson and Lindenstrauss shows that a set of n points in high dimensional Euclidean space can be mapped into an `O(log n/ϵ2)-dimensional` Euclidean space such that the distance between any two points changes by only a factor of `(1 ± ϵ)`. In this note, we prove this theorem using elementary probabilistic techniques.
*Dasgupta, Sanjoy, and Anupam Gupta. "An elementary proof of a theorem of Johnson and Lindenstrauss." Random Structures & Algorithms 22.1 (2003): 60-65.*