The Problem
Porch.com connects homeowners with local service professionals—electricians, plumbers, roofers, painters, and more. When a homeowner searches for help, the platform needs to rank professionals by relevance, availability, and proximity.
But here's the challenge: professionals rarely define their service areas explicitly. Some will drive 50 miles for a big job; others stay within a few ZIP codes. The data is sparse, noisy, and constantly changing.
Cold Start
New professionals have no job history. How do we estimate where they'll work?
Sparse Data
Most professionals have only a few completed jobs. Not enough to infer boundaries.
Implicit Behavior
Service areas aren't stated—they're revealed through patterns of accepted and declined jobs.
The Approach
We combined multiple unsupervised learning techniques to build a service area scoring system that improves with each completed job.
Geographic Priors from CBSAs
Core Based Statistical Areas (CBSAs) are Census-defined metropolitan and micropolitan regions. They provide natural boundaries for service area estimation.
- Metropolitan areas cluster population centers
- Professionals typically serve within their CBSA
- CBSAs capture commuting patterns and economic regions
We used CBSA membership as a prior probability—professionals are more likely to serve locations within their home CBSA than outside it.
Learning from Job Locations
Historical job locations reveal actual service patterns. We applied spatial clustering to discover natural service boundaries.
- Custom density-based clustering to identify dense regions of completed jobs
- K-means for professionals with multiple service zones
- Kernel density estimation for smooth probability surfaces
Job locations are positive examples. We also inferred negative examples from declined leads and "out of area" responses.
Collaborative Filtering from Colleagues
The key insight: professionals at the same company often share service area patterns. A new electrician at "ABC Electric" likely serves the same regions as their colleagues.
- Matrix factorization to learn latent service area embeddings
- Propagate knowledge from data-rich to data-sparse professionals
- Company affiliation, specialty, and location as grouping features
This solved cold start: new professionals inherit service area estimates from similar colleagues.
Technical Methods
Unsupervised Clustering for Service Boundaries
We experimented with multiple clustering approaches to discover natural service area boundaries from job location data:
K-Means Clustering
Partitions locations into K clusters. Works well when professionals have distinct, non-overlapping service zones (e.g., "downtown" vs. "suburbs").
minimize Σ ||x - μ_k||² Tradeoff: Requires choosing K upfront; assumes spherical clusters.
Custom Density-Based Clustering
I developed a density-based approach from first principles—later discovering it was conceptually similar to DBSCAN. It identified dense regions of job locations while naturally filtering noise.
neighborhood density ≥ threshold Lesson: Sometimes you reinvent the wheel—but you deeply understand the axle.
Gaussian Mixture Models
Probabilistic clustering that models service areas as overlapping Gaussian distributions. Provides soft membership scores.
P(x) = Σ π_k N(x|μ_k, Σ_k) Tradeoff: Smooth probabilities are useful for ranking; assumes Gaussian shape.
Collaborative Filtering for Cold Start
We modeled the professional-location relationship as a sparse matrix and applied matrix factorization to learn latent representations:
By learning embeddings for both professionals and locations, we can predict service area scores for any professional-location pair—even for new professionals with no history. Professionals with similar colleagues, specialties, or home locations will have similar embeddings.
Impact
Improved ranking of professionals by geographic fit, reducing out-of-area leads
New professionals received relevant leads immediately via collaborative filtering
Higher acceptance rates by matching homeowners with professionals who actually serve their area
Go Deeper
Explore the foundational techniques used in this project.
Clustering Algorithms
Comprehensive guide to K-Means, DBSCAN, Gaussian Mixtures, and other clustering methods.
Unsupervised LearningMatrix Factorization
Learn how matrix factorization powers collaborative filtering in recommendation systems.
Collaborative FilteringAbout CBSAs
Understanding Core Based Statistical Areas and how they define metropolitan regions.
Geographic DataThe Cold Start Problem
Overview of cold start challenges in recommender systems and common solutions.
Recommendations