Pinnaya Research
Methodology
How 262,502 Reddit posts and 2,996,169 comments were collected, cleaned, classified, audited and tested, and what the resulting numbers can and cannot support.
- Companion to
- The Indian Dating Conversation
- Published
- 19 August 2026
- Analysis period
- 3 Aug 2024 to 3 Aug 2026
Disclosure of interest. Pinnaya published this study and competes in the market it examines. This page exists so that the findings can be checked, replicated in principle, and disbelieved on the evidence rather than on suspicion.
Check the work yourself
Free · CC BY 4.0 · No form
- For journalistsPress fact sheetPDF, 3 pagesDownload PDF
- For publishingChart packZIP, 4 PNGsDownload ZIP
- For analystsAggregate datasetCSV, 511 rowsDownload CSV
Reuse with attribution to Pinnaya.
Data provenance
The underlying records were obtained from the Arctic Shift public research archive, a free, open-source, volunteer-maintained archive of historical Reddit data built for researchers. They were not obtained from Reddit’s own data API, and Reddit’s servers were not queried for data at any point. Requests went to the archive’s public posts and comments search endpoints, rate-limited and retried on throttling.
This study redistributes none of that source material. What is published is 511 derived aggregate statistics: counts, shares, confidence intervals and quarterly series. No post text, no comment text, no quotation, no permalink, no post or comment identifier and no author identifier appears in this study or in its dataset. Author identifiers were never collected at any stage.
Collection design
The corpus covers 48 Reddit communities across three collection strata: eight India-focused dating, relationship and matrimony communities traversed exhaustively; 31 India-national and Indian-city communities reached through 1,678 keyword search tasks; and nine global dating communities searched for India-related terms. Collection ran over the window 3 August 2024 to 3 August 2026 inclusive of the start and exclusive of the end.
| Stratum | Communities | Tasks | Unique posts | Comment coverage |
|---|---|---|---|---|
| Exhaustive community sweep | 8 | 11 | 202,621 | Complete |
| India and city community keyword search | 31 | 1,678 | 56,156 | Partial |
| Global dating community, India keyword | 9 | 36 | 3,725 | Partial |
Three additional community sweeps returned zero posts and are absent from the corpus. Comments were collected exhaustively for the eight swept communities and for 13,800 additional posts elsewhere.
Deduplication, joining and quality gates
265,311 raw post rows deduplicate to 262,502 unique posts, with 2,809 duplicates produced by overlapping discovery queries. 2,996,169 comment rows contained zero duplicates, verified by extracting the identifier of every row. All 2,996,169 comments joined successfully to a captured post. No records fell outside the collection window.
Comments were then filtered for analysability. 168,581 bot and automoderator comments, 3,001 moderator-boilerplate comments and 302,975 low-information comments, defined as fewer than three words or a pure reaction token, were separated out before any text analysis. 2,521,779 comments, 84.2 percent of the total, remain. Repeated-template detection ran on a rolling hash of the first 300 normalised characters with a threshold of 25 occurrences.
Post body text is present and usable in 155,426 posts, deleted or removed in 73,012, and empty in 34,064, the last group being image and link posts. Encoding damage is negligible at four posts and three comments. Posts flagged not-safe-for-work number 7,677, or 2.9 percent; they remain in quantitative counts and were excluded from every qualitative output.
Classification
Classification is rule-based, not model-based. Each post receives a relevance tier and tags across six dimensions: journey stage, primary need, friction, outcome, platform or channel, and explicitly-stated context segment. Every tag traces to a published pattern. Demographic attributes are recorded only where the author states them.
Matching runs in two stages: a cheap literal-substring gate over per-category trigger sets, then a precise pattern applied only to categories whose triggers are present. The gate was verified to lose nothing against a single-pass reference matcher over 6,400 document-dimension pairs, so it is a speed optimisation with no effect on results.
Relevance tiers
- High. A post in one of the eight dedicated communities carrying any domain signal, or a post elsewhere carrying a strong domain signal such as a named platform or matrimonial vocabulary in an India context, or at least two general domain signals in an India context.
- Medium. Weaker or partial signal. Used only in sensitivity testing, never in a headline figure.
- Low. No domain signal, or a false-positive pattern such as a cricket match, a date of birth or a business partner with nothing to offset it.
The analytical base is high-relevance posts of at least 15 words: 127,140 posts. Everything reported as a headline figure uses that base and states it.
Segment rules. Gender is taken only from explicit self-reference, such as an age-and-gender token or a first-person statement. City, age band, orientation, student or working status, childfree preference and marital history follow the same rule. Nothing is inferred from writing style, username, topic or stereotype. This is why gender is available for only 51.3 percent of the base and city for only 13.9 percent.
Validation, and the two corrections it forced
Two stratified manual audits were run against the raw text. Precision on the high-relevance tier is 92.8 percent, with a Wilson 95 percent confidence interval of 84.0 to 96.9 percent. Precision differs sharply by stratum: 98.0 percent inside the eight exhaustively-collected communities and 80.0 percent in the keyword-discovered ones. Every headline result is therefore also reported for the exhaustive subset.
| Tier | Audited | Verdict | Precision |
|---|---|---|---|
| High, all strata | 69 | 64 correct, 2 borderline, 3 false positives | 92.8% |
| High, exhaustive sweep | 50 | 49 correct | 98.0% |
| High, keyword-discovered | 15 | 12 correct | 80.0% |
| High, global dating communities | 4 | 4 correct | 100% |
| Medium | 48 | 36 on-topic but low information | 75.0% |
| Low | 34 | 32 correct exclusions | 94.1% |
Ambiguity and borderline rate across the sample: 3.3 percent. Samples were drawn with fixed random seeds so they can be regenerated.
The first audit round found two systematic errors, both corrected before any result in this study was calculated:
- India context was under-detected. Posts in Indian city and national communities were only treated as India-context when the body text contained an India-related word. Those communities are India-context by construction. Correcting this moved 11,655 posts from low to high relevance, 7,769 from low to medium, and 17,784 from medium to high.
- A false-positive override was too aggressive. Generic exclusion patterns such as date of birth and business partner were being applied inside the eight dedicated dating communities, demoting genuinely on-topic posts. Core-community posts are now scored on domain signal alone. 45 posts were corrected.
Publishing the errors is deliberate. A classification pipeline that never needed correcting was never audited.
Statistical standards
The unit of analysis is the post. Comments are used for engagement and context only and never as independent respondents, because comments are nested within posts and within communities. Every proportion carries a Wilson 95 percent confidence interval. Group differences are reported only when intervals do not overlap. A minimum cell size of 300 applies to every segment contrast.
- Cluster bootstrap. Comments-per-post figures use 200 bootstrap draws resampling posts, not comments.
- Trend rule. A change is called a trend only when it is sustained, interval-separated, and still present when the analysis is restricted to the eight exhaustively-collected communities. Two signals passed: dating fatigue and burnout language. One failed the second test and is published as a negative result.
- Two views. Corpus view is the dataset as collected. Balanced view is the mean of within-community shares across the eight exhaustively-collected communities, giving each equal weight. Both are published.
- No causal claims. Co-occurrence of themes is not mechanism, and nothing in this study is presented as cause.
- No population extrapolation. Findings describe the captured conversation.
| Scenario | Base size |
|---|---|
| A. Headline base: high relevance, at least 15 words | 127,140 |
| B. Widened to include medium relevance | 163,548 |
| C. Restricted to the eight exhaustively-collected communities | 86,846 |
| D. Largest single community removed | 88,904 |
| E. Leave-one-community-out, top eight run individually | varies |
| F. Posts with no collected comments excluded | 92,348 |
Maximum absolute shift in any individual friction under leave-one-community-out: 2.24 percentage points. Rank order of the leading frictions is stable across all scenarios.
Privacy and publication rules
No individual post, comment, quotation, permalink or author identifier is published anywhere in this study or its dataset. Only aggregate statistics are released. Author identifiers were never collected in the first place, and no attempt was made to identify anyone at any stage.
Internal analysis excluded from any qualitative handling: posts flagged not-safe-for-work, posts by apparent minors identified through self-stated age or school-grade references, and explicit sexual-violence, self-harm and abuse narratives. Those exclusion rules were applied to internal working material even though no qualitative material is published, because the rule should not depend on the publication decision.
Limitations
- Reddit is not India. The corpus skews young, urban, English-writing and internet-native. No figure is a population estimate.
- The strata behave differently. 98.0 percent audited precision in the exhaustive sweep against 80.0 percent in keyword discovery.
- Comment coverage is uneven. 71.9 percent of posts have collected comments; coverage is exhaustive only in eight communities. Comments-per-post is uninterpretable elsewhere and is flagged where shown.
- Segments are self-reported and unverified. Gender for 51.3 percent of the base, city for 13.9 percent.
- Rule-based classification misses nuance. Sarcasm, irony and code-mixed Hindi, Malayalam and Hinglish are systematic error sources. Communities outside the metros and outside English are the most affected.
- Removed content is invisible. Deleted and removed material was not retained at collection, so safety figures are floors.
- Disclosure is not experience. A group difference in how often a friction is named may reflect willingness to name it.
- Scores are point-in-time. Post and comment scores are values at retrieval, not final values.
- Journey stage uses a convention. Most-advanced-stage-wins is defensible but not unique and favours later stages in multi-stage posts.
- The final quarter is partial. 2026 Q3 covers one month and 7,079 posts.
- One post lost. A single post identifier, 0.0004 percent of the corpus, was lost during a processing checkpoint repair. It is disclosed rather than quietly reconciled.
Reuse
The aggregate dataset of 511 derived statistics is published under CC BY 4.0. Attribute to Pinnaya. Underlying Reddit content is not redistributed and is not available on request.
About Pinnaya
Pinnaya is an intentional dating app for urban Indian professionals. Every member is verified with government ID and face matching, women’s profiles stay hidden by default until they choose to reveal them, and each user holds a maximum of three active matches at a time. Pinnaya is positioned between swipe-based dating apps and traditional matrimonial platforms.
Pinnaya is live on iOS, Android and at Pinnaya.com. Pinnaya is incubated at IIMA Ventures and has filed two provisional patents. Pinnaya published this study and competes in the market it examines.
Questions and corrections
Write to care@pinnaya.com. Methodology challenges are welcome and are the point of publishing this page. If you find an error we will publish the correction on the study page with a dated note rather than editing quietly.
Methodology companion to The Indian Dating Conversation. Published 19 August 2026.