CIS 5200 · Machine Learning · University of Pennsylvania

Community Detection on Heterogeneous Academic Networks

A 40M-edge academic graph reduced by 76% without losing its semantics, then embedded with Metapath2Vec and HGT — reaching 0.91 accuracy and driving a reviewer-matching and community-discovery tool.

Community Detection on Heterogeneous Academic Networks

Period

Fall 2025

Stack

PyTorch · Graph Neural Networks · HGT · HAN · Metapath2Vec · Streamlit

Highlights

  • 010.91 classification accuracy with HGT on Metapath2Vec features — against 0.44 on random features, which is the whole finding
  • 02Graph reduced from 37,791 nodes and 40.3M edges to 22,741 nodes and 9.7M edges: 76% fewer edges, and HAN/HGT became trainable at all
  • 03Reduction validated three ways: degree distributions and edge-type ratios preserved, Jensen-Shannon divergence of label distributions ≈ 0, and 70% nearest-neighbour overlap in embedding space
  • 04Benchmarked Metapath2Vec (NMI 0.502 / ARI 0.463 / acc 0.394), HAN (0.706 / 0.762 / 0.88) and HGT (0.702 / 0.759 / 0.91)
  • 05Shipped two Streamlit applications on the learned embeddings: reviewer matching with co-author exclusion, and a community explorer that ranks sub-communities by cohesion

What it drives

Reviewer matching. Pick a paper and an embedding backend, and the tool ranks candidate reviewers by embedding similarity — optionally excluding co-authors, which is the conflict-of-interest constraint that makes the problem worth automating. The label column shows how many of the top-K actually match the paper's field.
Reviewer matching. Pick a paper and an embedding backend, and the tool ranks candidate reviewers by embedding similarity — optionally excluding co-authors, which is the conflict-of-interest constraint that makes the problem worth automating. The label column shows how many of the top-K actually match the paper's field.
The community explorer. Within a chosen field it clusters authors into sub-communities, ranks them by cohesion, surfaces the core authors of each, and computes a similarity ranking around a focal author.
The community explorer. Within a chosen field it clusters authors into sub-communities, ranks them by cohesion, surfaces the core authors of each, and computes a similarity ranking around a focal author.
t-SNE of the HGT embeddings, coloured by the four author labels (DB, DM, AI, IR). The fields separate cleanly, which is the visual form of the 0.91 accuracy — and the scattered off-colour points are the genuinely cross-disciplinary authors.
t-SNE of the HGT embeddings, coloured by the four author labels (DB, DM, AI, IR). The fields separate cleanly, which is the visual form of the 0.91 accuracy — and the scattered off-colour points are the genuinely cross-disciplinary authors.

Write-up

Academic networks are not one graph. They are authors, papers, terms and venues tangled together by relations that mean different things, and collapsing that into a homogeneous graph throws away exactly the structure that defines a research community. The dataset — the heterogeneous DBLP network from Ji et al. and Gao et al. — has roughly 14k authors, 14k papers, 8.9k terms and 20 conferences, with four field labels: databases, data mining, AI and information retrieval.

The first problem was purely practical: expanded to its full relational form the graph has about 40 million edges, and neither HAN nor HGT will train on that. Sampling it down naively would have been easy and wrong — it would change the thing being measured. So the reduction keeps every labelled paper, keeps every author attached to one, stratifies the sampling of remaining papers by degree, then pulls in the terms and conferences those papers touch and rebuilds a consistent subgraph. That took it to 22,741 nodes and 9.7M edges — 76% fewer edges, with all 20 conferences intact.

Then I had to show the reduction had not quietly destroyed the semantics, which needed three separate checks rather than one: degree distributions and edge-type ratios still match, the label distribution has a Jensen-Shannon divergence of about zero against the original, and 70% of nearest neighbours in embedding space survive the cut. Structure, labels, and learned geometry — a reduction could pass any one of those and still be useless.

On models, the interesting result is not which architecture won but how much the features mattered. On random input features HGT scores 0.44 accuracy and a poor NMI of 0.327 — worse-clustered than HAN. Give both models Metapath2Vec embeddings as input and HGT jumps to 0.91 and HAN to 0.88. The type-aware attention has plenty of capacity, but it cannot invent semantics that were never in the features; the metapath walks supply exactly that. Worth saying plainly: on the clustering metrics HAN is fractionally ahead (NMI 0.706 vs 0.702, ARI 0.762 vs 0.759). HGT wins on accuracy and on how consistent its core-author rankings are, which is why it became the final model, but it does not dominate on every measure.

The last piece is what the embeddings are actually for. Two Streamlit tools sit on top: reviewer matching, which ranks candidate reviewers for a paper by similarity while excluding co-authors and checking field labels, and a community explorer that clusters authors within a field, ranks the sub-communities by cohesion and identifies the core researchers in each. Those are the concrete forms of the stakeholder problems the project started from — conferences needing conflict-free reviewer assignment, and universities and funding agencies needing to see where cohesive clusters and emerging subfields actually are.

Team project with Lakshay Naresh Ramchandani and Ishita Rai.

Links