A behavioral and machine-learning analysis of the Su & Hu online-dating dataset.
Dataset source: Su, X., Hu, H. et al. Gender-specific preference in online dating. EPJ Data Science 8, 12 (2019). https://doi.org/10.1140/epjds/s13688-019-0192-x — https://link.springer.com/article/10.1140/epjds/s13688-019-0192-x
This project asks a deliberately concrete question:
What observable characteristics of a male dating profile are associated with female attention and active engagement, which combinations add information beyond individual features, and how much of that behavior can be predicted out of sample?
The analysis uses observed recommendation, click, and message events rather than stated survey preferences. The distinction is important. A recommendation creates an opportunity for attention, a click indicates that the female user opened the male profile for more information, and a message is a stronger observed engagement action.
The project therefore separates two behavioral stages:
RECOMMENDED
|
v
PROFILE CLICK
|
v
MESSAGE
This separation also controls for feature visibility. The source documentation states that the recommendation stage exposed only a limited subset of the candidate profile, while a click exposed the detailed profile. That means education, income, occupation, marital status, and similar attributes cannot simply be treated as if they were visible at the first decision point.
The resulting analysis is a layered investigation:
RAW USER PROFILES + BEHAVIOR LOG
|
v
FEMALE -> MALE EXPOSURES
|
+-------+--------+
| | |
v v v
AGE HEIGHT PROFILE
| | ATTRIBUTES
+-------+--------+
|
v
INDIVIDUAL EFFECTS
|
v
PAIRWISE INTERACTIONS
|
v
FEATURE GROUPS
|
v
NONLINEAR PREDICTION
|
v
FINAL INTERPRETATION
The executed analysis reports:
| Finding | Result |
|---|---|
| Female profiles | 203,843 |
| Male profiles | 344,552 |
| Raw behavior records | 8,599,012 |
| Unique female to male exposure pairs | 1,654,502 |
| Pair click rate | 3.285% |
| Pair message rate | 0.567% |
| Strongest individual feature by permutation importance | Candidate age |
| Largest single-feature click-rate spread | Candidate age, about 7.38 percentage points |
| Best reduced feature family | Physical + socioeconomic |
| Best reduced ROC-AUC | 0.627 |
| Full nonlinear model ROC-AUC | 0.641 |
| Full nonlinear model PR-AUC | 0.374 |
| Strongest compatibility signal | Age preference fit |
| Click rate inside age requirement | 3.484% |
| Click rate outside age requirement | 2.124% |
The central result is therefore not "women prefer feature X".
The stronger statement is:
Candidate age is the strongest individual predictive signal in this dataset, while combining physical/age information with socioeconomic information produces the strongest reduced feature family. Among explicit compatibility variables, satisfying the viewer's age requirement produces the largest observed engagement difference.
These are associations in a historical online-dating dataset. They are not universal laws and they are not causal estimates.
The primary research question is:
What observable characteristics of a male profile are associated with female attention and active engagement, and which combinations of characteristics provide additional explanatory power beyond individual features?
The study answers this through four progressively more demanding questions.
Examples include:
- age
- height
- education
- income
- occupation
- marital status
- housing
- car ownership
- privacy
- avatar state
- credit rating
Examples include:
age × education
age × income
height × age
education × income
education × occupation
marital status × age
compatibility × socioeconomic status
The project compares interpretable feature families rather than only ranking raw columns.
The main predictive split is grouped by female user. A woman's records are kept out of the training data when she appears in the evaluation set.
This makes the evaluation closer to the actual question:
Can the learned relationship generalize to women the model did not train on?
Dataset source & citation: Su, X., Hu, H. et al. Gender-specific preference in online dating. EPJ Data Science 8, 12 (2019). DOI: 10.1140/epjds/s13688-019-0192-x — Publisher page: https://link.springer.com/article/10.1140/epjds/s13688-019-0192-x. The
data/folder in this repository (profile_f.txt,profile_m.txt,matching_data.txt) is derived from the dataset released with that publication.
The source documentation describes three tables:
profile_f.txt
profile_m.txt
matching_data.txt
The two profile tables contain 35 attributes. The behavior table contains:
USER_ID_A
USER_ID_B
ROUND
ACTION
where ACTION can be:
rec
click
msg
The recommendation event is the first observable stage. The user can then open the recommended profile, producing click, and can subsequently contact the other user, producing msg.
A pair can be recommended more than once, and a click or message can occur more than once. The recommendation round preserves temporal ordering between recommendation rounds.
That gives the project an unusually useful structure:
FEMALE VIEWER
|
| recommendation round
v
MALE CANDIDATE
|
+-----------+-----------+
| |
click no click
|
message
The analysis converts the event stream into a compact pair-level exposure representation so repeated raw events do not have to remain in every downstream table.
The supplied source contains:
- 203,843 female profiles
- 344,552 male profiles
- 8,599,012 historical behavior records
- 1,654,502 unique female to male exposure pairs
The observed funnel is highly asymmetric:
Recommendation
████████████████████████████████████████
Click
█
Message
▏
The pair-level rates are approximately:
- click: 3.285%
- message: 0.567%
This means that "engagement" is not a common event. The model therefore needs to work in an imbalanced setting and should not be evaluated through accuracy alone.
The pipeline figure summarizes the complete path from source tables to behavioral outcomes, feature construction, group analysis, prediction, and final interpretation.
The important design choice is that the analysis hierarchy is explicit. Individual variables are not immediately fed into a black box. They are first examined descriptively, then paired, then grouped, and only after that used in predictive models.
This is the conceptual hierarchy behind the project:
A feature can look important alone but add almost nothing once another correlated feature is present. The reverse can also happen: a weak individual effect can become useful because it interacts with another feature.
The funnel is the most important behavioral framing in the repository.
A recommendation is an exposure event. A click means the user chose to inspect the candidate in more detail. A message is a stronger active action.
Therefore:
recommendation
!=
attention
attention
!=
active outreach
This is why the project keeps click and message analysis separate.
The pair-level data produce:
Click rate = 3.285%
Message rate = 0.567%
The message rate is substantially lower, which means that many attributes may be useful for attracting initial attention without necessarily being sufficient to produce active outreach.
The source platform did not expose every male attribute at the recommendation stage.
At recommendation time, the viewer could see a limited representation of the candidate. Detailed profile attributes became available after the click.
This creates a strict modeling rule.
Appropriate variables include:
- candidate age
- region
- sub-region where applicable
- avatar state
- viewer age
- age gap
- location compatibility
- age preference fit
- other attributes that are demonstrably available before profile inspection
The analysis can additionally use detailed profile variables such as:
- education
- income
- occupation
- marital status
- children
- housing
- car ownership
- privacy
- credit rating
- religion
- ethnicity
- detailed preference fields
This prevents the most obvious form of temporal leakage.
A model cannot legitimately explain a click using a variable the viewer had not yet been shown.
The first substantive question is simple:
What happens when we look at one feature at a time?
The notebook evaluates each major attribute using observed click and message rates.
Candidate age is the strongest individual signal in the project.
It has:
- the largest permutation importance in the nonlinear model
- the largest raw click-rate spread
- an observed spread of approximately 7.38 percentage points across the analyzed age bins
This is substantially larger than the raw differences associated with several socioeconomic and profile variables.
The result does not mean "there is one universally attractive age". It means that the distribution of female engagement varies most strongly with candidate age in this dataset.
The exact relationship is nonlinear, which is why the binned response plot is more informative than a single linear coefficient.
Height has a visible relationship with engagement, but its overall spread is smaller than age.
This is an important result because it prevents the analysis from turning into the simplistic claim that a physical attribute dominates everything else.
Height contributes signal, but it is not the strongest single feature.
Education has a much smaller raw click-rate spread than age.
That does not make education irrelevant.
The more important question is whether education becomes useful:
alone
versus:
education + income
education + occupation
education + age
That is why education appears again in the interaction analysis.
Income shows an observable association with engagement, but its univariate spread is weaker than the age signal.
Income becomes more useful when treated as part of the broader socioeconomic block.
This is exactly the type of variable that can be undervalued by a purely univariate analysis and overvalued by a simplistic interpretation of raw group differences.
Marital status has one of the larger individual engagement differences after age.
It also interacts naturally with age because relationship state can carry different meanings at different candidate ages.
Therefore the project does not treat marital status as an isolated ranking variable.
Occupation contains categorical information that is difficult to summarize through a single ordered number.
The analysis therefore keeps occupation categorical and evaluates its predictive value inside the socioeconomic group rather than pretending that occupation codes have a meaningful linear distance.
Avatar presence is interesting because it connects to the visibility and presentation stage.
However, its raw engagement spread is much smaller than the age effect.
This is another useful negative result: an attribute can be intuitively important without being the strongest observed signal in the actual behavioral data.
The individual-feature summary provides the clearest high-level ranking.
Candidate age
The analysis shows progressively smaller raw spreads for:
- marital status
- housing
- occupation
- height
- privacy
- income
- credit level
- education
- avatar state
The distinction between "raw engagement spread" and "unique predictive contribution" is essential. A large univariate gap does not automatically mean that the feature is uniquely informative once other variables are included.
The dataset includes explicit mate-preference fields for the female viewer.
That makes it possible to construct compatibility indicators instead of only looking at the man's attributes.
The strongest explicit preference-fit signal is age compatibility.
A candidate who falls inside the female viewer's stated age requirement has:
3.484% click rate
compared with:
2.124% click rate
for candidates outside that requirement.
That is an absolute difference of about:
1.36 percentage points
and a relative increase of roughly:
64%
in the observed click rate.
This is one of the strongest results in the entire project because it directly links:
female stated preference
+
male candidate attribute
=
observed behavior
The age-gap analysis shows that compatibility is not adequately described by a single candidate-age variable.
Two men of the same age can receive different engagement from women of different ages because the viewer's own age changes the meaning of the age gap.
This is why the strongest model uses both candidate-side and viewer-side information.
Height preference fit is informative but substantially weaker than age preference fit in the aggregate results.
The implication is not that height is irrelevant. It means that the stated age requirement provides a cleaner and stronger preference-matching signal in this dataset.
The project deliberately moves beyond additive feature importance.
The interaction analysis asks:
Does one feature change the meaning of another feature?
Examples include:
age × education
age × income
height × age
education × income
education × occupation
marital status × age
compatibility × socioeconomic status
A positive interaction means that the joint configuration is more associated with engagement than would be expected from simply adding two independent feature effects.
A negative interaction means the joint configuration contributes less than the additive expectation.
Suppose education has only a modest individual association.
That does not imply:
education does not matter
It could instead mean:
education matters differently for different income levels
education matters differently at different ages
education matters differently by occupation
The pairwise analysis exists specifically to distinguish these cases.
The analysis groups individual variables into interpretable families.
| Feature group | Main variables |
|---|---|
| Physical and age | age, height, age gap |
| Socioeconomic | education, income, occupation, credit |
| Lifestyle and assets | housing, car ownership |
| Relationship state | marriage, children |
| Visibility and presentation | avatar, privacy |
| Compatibility | age fit, height fit, education fit, region fit |
The goal of grouping is not to manufacture a single "attractiveness score".
It is to ask:
Which category of information contributes the most useful predictive signal when considered together?
The reduced feature-group experiment produces the following ordering:
1. Physical + socioeconomic
2. Physical + age
3. Other combinations containing lifestyle / relationship information
4. Compatibility alone
Physical + socioeconomic
This is the strongest reduced group in the current experiment, with:
ROC-AUC = 0.627
The full nonlinear feature set reaches:
ROC-AUC = 0.641
That gives a useful interpretation.
The strongest group already captures most of the model's ranking ability.
The remaining variables and nonlinear interactions improve the result further, but not dramatically.
This is the main result the project was designed to answer.
It has:
- the strongest permutation importance
- the largest observed click-rate spread
- the clearest evidence of nonlinear response behavior
Candidates within the viewer's stated age requirement receive:
3.484% click rate
versus:
2.124%
outside the stated requirement.
This is the strongest aggregate preference-fit result.
Containing:
Physical / age
+
Education
Income
Occupation
Credit
This family reaches:
ROC-AUC = 0.627
on the reduced group experiment.
The full model reaches:
- ROC-AUC = 0.641
- PR-AUC = 0.374
on the held-out female-user evaluation.
The complete model is therefore better than the strongest reduced group, but the improvement is moderate rather than enormous.
That is an important result.
It suggests that the main explanatory structure is already concentrated in a relatively small set of broad dimensions:
AGE / PHYSICAL
+
SOCIOECONOMIC
+
INTERACTIONS
rather than requiring every profile attribute to be equally important.
This question requires a precise distinction.
There is no scientifically justified universal raw-feature threshold such as:
height >= X
income >= Y
education >= Z
that can be claimed from this analysis.
The observed relationships are probabilistic and nonlinear.
Likewise, the feature-group analysis identifies which groups add predictive information, not a single scalar threshold at which a group suddenly becomes "enough".
The correct threshold is a model decision threshold, applied to the predicted probability after the model has been trained and calibrated.
For example:
P(click) >= threshold
|
v
predicted engagement
The threshold should be selected on validation data according to the operating objective:
- maximize F1
- prioritize recall
- prioritize precision
- control false positives
- or optimize a business-specific cost
The current measured results establish the strongest feature and group structure, but they do not provide a defensible universal raw-value threshold for the physical + socioeconomic feature family. The project therefore does not invent one.
The evidence supports:
on the held-out female-user evaluation.
The additional performance over the strongest reduced group is modest:
0.641 - 0.627 = 0.014 ROC-AUC
That means the compact physical + socioeconomic representation already captures a large share of the predictive structure found by the full model.
The results should not be rewritten as:
Women like older men.
or:
Women like rich men.
or:
Women prefer tall men.
Those claims are too strong for the evidence.
The dataset measures behavior on one historical platform, under one recommendation system, with one particular population and a limited representation of the candidate profile.
The correct interpretation is:
Within this dataset, candidate age contains the strongest individual predictive signal for female click behavior. Physical and socioeconomic information form the strongest reduced predictive family, and explicit age compatibility has the strongest aggregate preference-fit association.
The distinction between prediction and causation remains important throughout.
A feature can be useful for prediction because it is correlated with another unobserved attribute.
The most interesting conclusion is not that age wins.
It is that the model does not need every variable equally.
The structure looks closer to:
FEMALE CLICK BEHAVIOR
|
+--------------+--------------+
| |
v v
CANDIDATE AGE PHYSICAL / AGE
|
v
SOCIOECONOMIC
|
v
INTERACTIONS
|
v
FINAL PREDICTION
The feature-group experiment shows that a compact block of observable information already carries most of the useful predictive signal.
This suggests a useful next research direction:
Can the same feature hierarchy reproduce across other dating platforms, cultures, years, and recommendation systems?
That is the correct next step for testing generalization.
The notebook is organized as a research walkthrough:
01 Research framing
02 Environment and paths
03 Source data verification
04 Data dictionary
05 Profile quality audit
06 Interaction log audit
07 Female -> male exposure construction
08 Behavioral funnel
09 Candidate-side descriptive analysis
10 Viewer-side preference analysis
11 Compatibility features
12 Individual feature analysis
13 Pairwise feature interactions
14 Feature-group ablation
15 Predictive modeling
16 Model diagnostics
17 Feature importance
18 Results synthesis
19 Conclusions
20 Limitations and extensions
Each section is designed to answer one analytical question rather than simply execute code.
The raw behavior file contains millions of records, so the notebook avoids repeatedly constructing a giant in-memory representation.
The processing strategy is:
raw interaction file
|
v
chunked reads
|
v
female -> male filtering
|
v
within-chunk aggregation
|
v
compact pair-level state
|
v
descriptive analysis
|
v
controlled modeling sample
The descriptive statistics are derived from the complete filtered exposure table.
The predictive model uses a controlled sample with a fixed seed and grouped evaluation by female user.
This keeps memory requirements reasonable on a standard laptop without changing the analytical logic.
The figures below are embedded directly from assets/ so the README functions as a visual research report rather than a text-only index.
The repository contains a coherent visual set.
assets/research_pipeline.pngassets/analysis_hierarchy.pngassets/feature_visibility.png
assets/behavioral_funnel.pngassets/behavioral_funnel_counts.png
assets/male_age_click_rate.pngassets/male_height_click_rate.pngassets/male_education_click_rate.pngassets/male_income_click_rate.pngassets/male_marriage_click_rate.pngassets/occupation_click_rate.pngassets/male_avatar_click_rate.png
assets/preference_fit.pngassets/age_gap_response.pngassets/height_gap_response.pngassets/education_income_interaction.png
assets/individual_feature_effects.pngassets/feature_group_ablation.pngassets/permutation_importance.pngassets/model_comparison.pngassets/confusion_matrix.png
assets/conclusion_dashboard.png
The README deliberately uses the figures beside the claims they support so that the results can be inspected without opening the notebook.
The source dictionary defines 35 profile fields. The most important fields used in this project include:
| Field | Meaning |
|---|---|
Uid |
User ID |
sex |
Gender |
register_time |
Registration time |
last_login |
Last login time |
birth_year |
Birth year |
Birthday |
Date of birth |
work_location |
Region |
work_sublocation |
Sub-region |
status |
User status |
login_count |
Login frequency |
education |
Education level |
house |
Housing situation |
auto |
Car ownership |
marriage |
Marital status |
children |
Child situation |
industry |
Occupation |
privacy |
Photo viewing permissions |
level |
Credit rating |
nation |
Ethnicity |
height |
Height |
income |
Monthly income category |
avatar |
Avatar availability |
belief |
Religion |
match_min_age |
Minimum desired partner age |
match_max_age |
Maximum desired partner age |
match_min_height |
Minimum desired partner height |
match_max_height |
Maximum desired partner height |
match_certified |
Credit-rating preference |
match_marriage |
Marital-status preference |
match_education |
Education preference |
match_edu_more_than |
Whether education can exceed the selected requirement |
match_avatar |
Avatar requirement |
match_work_location |
Preferred region |
match_work_sublocation |
Preferred sub-region |
The full source dictionary remains the authoritative reference for categorical encodings.
The notebook is the primary executable artifact.
The first notebook cell installs required dependencies.
All project paths are resolved relative to the project root.
The notebook records:
- source filenames
- row counts
- random seed
- sampling parameters
- model configuration
- evaluation configuration
in the project output manifest.
No machine-specific absolute paths are required.
A click is not a direct measurement of physical attraction.
It is an observed action that can reflect interest, curiosity, relevance, or many other motivations.
A message is a stronger signal, but it still does not establish relationship success.
We only observe what the platform chose to recommend.
A man who is never recommended to a woman cannot appear as a negative interaction for that woman.
Therefore the model learns:
preference conditional on recommendation exposure
not:
universal preference over every possible man.
The dataset comes from a historical Chinese dating platform.
The results should not be interpreted as a current universal description of women in other countries or on present-day dating apps.
The results describe association and predictive structure.
They do not establish that changing a man's age, education, income, or height would cause a woman to behave differently.
The source provides an avatar indicator rather than a modern multimodal profile-photo representation.
Therefore the project cannot answer questions about:
- facial attractiveness
- smile
- clothing
- body composition
- photo composition
- travel photos
- group photos
- gym photographs
Those would require a different dataset or a richer profile corpus.
This project uses the Su & Hu online-dating dataset released with:
Su, X., Hu, H. et al. Gender-specific preference in online dating. EPJ Data Science 8, 12 (2019). https://doi.org/10.1140/epjds/s13688-019-0192-x Publisher: https://link.springer.com/article/10.1140/epjds/s13688-019-0192-x — Open Access, EPJ Data Science (SpringerOpen).
Please cite the original publication if you reuse the dataset. The dataset files in data/ (profile_f.txt.gz, profile_m.txt.gz, matching_data.txt.gz) and field dictionary correspond to the supplementary data described in that article.
The strongest answer supported by the current analysis is:
The practical conclusion is not a checklist of physical or socioeconomic thresholds.
It is a hierarchy:
- Candidate age is the strongest individual feature in this dataset.
- Physical plus socioeconomic information is the strongest reduced feature group.
- Age compatibility is the strongest explicit viewer preference-fit signal.
- Combining these features with nonlinear interactions produces the strongest predictive model.
- No universal raw-feature cutoff is justified by the evidence. The appropriate binary decision threshold belongs to the calibrated prediction system, not to any single feature group.
That is the result the data support.

















