Opposition to zoning reform frequently rests on an empirical claim: that allowing apartments, condominiums, or townhomes near single-family homes will reduce those homes' values. The claim is testable wherever multi-family and single-family housing already coexist at fine spatial grain. Oak Park, Illinois, an inner-ring Chicago suburb of about 4.7 square miles, is well suited to the test: nearly a thousand small multi-family buildings sit mid-block inside otherwise single-family residential fabric, a legacy of pre-1940s development, and roughly a third of the village's detached homes are within 200 feet of one.
This paper measures whether that proximity is capitalized into assessed market values. The identification problem is that multi-family buildings are not randomly located: they concentrate on arterial corridors whose traffic and commercial activity independently depress residential values. I address this with three devices: (i) an explicit classification of every multi-family building as corridor or embedded; (ii) distance-ring regressions in the style of Linden and Rockoff (2008) with a fixed effect for each multi-family building, so that identification comes only from comparing homes at different distances around the same building; and (iii) dose-response specifications that count nearby buildings rather than measuring distance to the nearest one.
All data originate from Cook County's public Socrata datasets:
assessed values (uzyt-m557), property characteristics
(x54s-btds), parcel addresses (3723-97qp),
address points (78yw-iddh), and parcel sales
(wvhk-k5uv), ingested into a SQLite snapshot; full URLs are
listed under Data Sources at the end of the paper.
Values are the 2026 mailed reassessment for Oak Park township; the assessor's
estimated market value is assessed value × 10 for class-2
residential property. The dependent variable is therefore the assessor's
modeled market value, not a transaction price, a limitation addressed
directly by the sale-price models of Section 8 and discussed in Section 10.
Outcome sample. All detached single-family homes (classes 202–209, 234, 278) with property characteristics, positive value, lot size between 500 and 50,000 sq ft, and coordinates: n = 9,596 of 9,648 class-eligible parcels (52 lacked coordinates; none lacked characteristics).
Exposure sample. 2,041 multi-family building locations: apartment and mixed-use buildings (classes 211, 212, 313–318, 391) located directly by address point; condominium buildings located by collapsing unit PINs to their site address and matching against the address-point layer via a deterministic cascade (89% of units located); and townhomes/row houses (classes 210, 295). Parcel records sharing an identical coordinate (condominium unit stacks, townhome rows geocoded to one address) are collapsed to a single building, so the 2,332 exposure parcel records reduce to 2,041 unique structures and dose counts refer to buildings rather than parcels. Homes' distances to buildings are measured address point to address point, so immediate adjacency typically registers as 50–150 feet. Condominium geocoding partly relies on nearest-street-number matching; its residual error is one reason the ring design below compares homes only within 800 feet of the same building.
Arterial exposure. Each home's own distance to the nearest address point on a named arterial street is computed alongside its multi-family distances, so the busy-street effect can be estimated directly on homes rather than inferred from building classification. Qualifying arm's-length sale transactions (2017–2025) are extracted for the outcome-validation models of Section 8.
A building is classified corridor if its address is on a named arterial street, if it lies within 150 feet of any address point fronting a named arterial (capturing corner buildings addressed on side streets), or if it lies within 300 feet of a commercial parcel (class 5xx). Otherwise it is embedded. The commercial threshold is swept over 200, 300, 400 feet in all specifications. At the 300-foot threshold, 1,152 buildings are corridor and 889 embedded.
The descriptive point of departure: homes within 100 feet of any multi-family building are assessed about 6% below homes 800+ feet away. The question is whether that reflects the buildings or their locations. The main specification is a hedonic ring regression:
ln Vi = Xiβ + Σk γk Bandik + μb(i) + εi
where Vi is home i's estimated market value; Xi is a full hedonic vector (log building and lot area, age and age squared, bedrooms, full and half baths, fireplaces, central air, garage capacity, basement type and finish, construction quality, repair condition, plus log distance to commercial and a within-200-ft-of-commercial indicator); Bandik indicates distance k ∈ {0–100, 100–200, 200–400} feet to the nearest embedded (separately, corridor) building, with 400–800 feet as the reference; and μb(i) is a fixed effect for home i's nearest multi-family building. The sample is restricted to homes within 800 feet of a building of the relevant type, so γk is identified only by comparing homes at different distances around the same building; the outer-ring homes on the same blocks are the controls. The design adapts the distance-ring logic of Linden and Rockoff (2008) to a cross-section: it removes between-building location differences, but unlike their event-study setting it has no before/after variation, so it bounds capitalization of existing proximity rather than identifying a treatment effect. Not every building's ring contains reference-band homes: at the 300-foot threshold, 155 of 543 embedded buildings have 400–800 ft homes (5,525 homes), and 128 have both a 0–100 ft home and a reference home; band coefficients are identified from those buildings, and a restricted specification drops the remainder entirely. Standard errors are clustered on the building. As robustness, the same bands are estimated in pooled models with no location fixed effects, 11 assessor-neighborhood fixed effects, and 1000-foot grid-cell fixed effects, clustered accordingly, and the key pooled model is re-inferred with Conley spatial-HAC standard errors.
Around embedded buildings the distance gradient is flat and statistically zero: homes within 100 feet of a mid-block multi-family building are assessed +0.9% (95% CI -0.2 to +2.0) relative to homes 400–800 feet from the same building; the confidence interval rules out discounts larger than 0.2%. Restricting to the 155 buildings whose rings actually contain reference-band homes (5,525 homes) gives +0.6% (t = 0.94), and adding each home's own arterial-distance controls gives -0.1%. Around corridor buildings a modest gradient exists (-3.2% within 100 feet, identified from 260 of 553 buildings with reference homes); for those buildings, distance to the building and distance to the arterial are the same measurement, so the estimate bounds street and building effects jointly.
| Band | Homes | Median value | Value/bldg sq ft | Med. bldg sq ft | Med. lot sq ft | Med. yr built |
|---|---|---|---|---|---|---|
| 0-100 ft | 937 | $570,000 | $322 | 1,836 | 5,676 | 1912 |
| 100-200 ft | 1,853 | $540,000 | $322 | 1,668 | 4,725 | 1914 |
| 200-400 ft | 3,062 | $550,000 | $325 | 1,752 | 5,128 | 1913 |
| 400-800 ft | 2,138 | $560,000 | $325 | 1,788 | 5,320 | 1916 |
| >=800 ft | 1,606 | $690,000 | $322 | 2,130 | 6,200 | 1927 |
Raw medians, no controls. Homes nearer embedded multi-family are older and smaller on smaller lots. This is the composition the hedonic controls absorb.
| Distance to nearest embedded MF | No location FE | Neighborhood FE | Grid FE (1000 ft) | Building FE (ring) |
|---|---|---|---|---|
| 0-100 ft | +0.6 (0.5) | +1.2 (0.8) | +1.1 (0.9) | +0.9 (0.5) |
| 100-200 ft | +0.5 (0.4) | +0.9 (0.5) | +0.5 (0.9) | +0.2 (0.5) |
| 200-400 ft | +0.9* (0.4) | +1.3* (0.5) | +0.6 (0.9) | +0.2 (0.4) |
| 400-800 ft | +0.7 (0.4) | +1.4* (0.5) | +0.7 (0.7) | ref. |
| Distance to nearest corridor MF | No location FE | Neighborhood FE | Grid FE (1000 ft) | Building FE (ring) |
|---|---|---|---|---|
| 0-100 ft | -6.8* (0.8) | -5.9* (1.6) | -4.3* (1.2) | -3.2* (0.9) |
| 100-200 ft | -4.8* (0.5) | -3.4* (1.5) | -1.8* (0.8) | -1.9* (0.6) |
| 200-400 ft | -3.6* (0.4) | -2.5* (1.2) | -0.8 (0.8) | -1.0* (0.5) |
| 400-800 ft | -2.4* (0.3) | -1.3 (0.8) | -0.4 (0.5) | ref. |
Percent effects, exp(γ)−1; robust/clustered SEs in parentheses (percentage points); * = |t| ≥ 1.96. Columns 1–3: pooled models, reference ≥800 ft, both band sets jointly estimated, full hedonic and commercial controls. Column 4: ring design, reference 400–800 ft around the same building. Neighborhood-FE SEs rest on 11 clusters and warrant caution; grid and building FE columns are the better-powered clustered estimates.
The corridor/embedded classification attributes the corridor discount to busy streets. That attribution is testable: each home's own distance to the nearest arterial is in the data. Regressing value on arterial-distance bands alone, with full hedonic and commercial controls, grid-cell fixed effects, and no multi-family terms of any kind, homes within 100 feet of an arterial are assessed -3.7% (t = -3.98) and homes at 100–200 feet -1.7%, relative to homes 800+ feet away. Busy streets carry their own discount, apartment or no apartment.
Estimating the multi-family bands and the home's own arterial bands jointly then decomposes the corridor discount: the corridor 0–100 ft coefficient falls from -4.3% to -2.5% (t = -2.17), with nothing beyond 100 feet, while the arterial gradient persists (-3.3% within 100 feet). Embedded remains zero (+0.5%). Most of the corridor discount is the street itself; what survives within 100 feet may reflect corridor buildings, address-point measurement error, or unmeasured frontage traits, and the design cannot separate those.
Nearest-building distance ignores intensity. If multi-family density itself were the disamenity, homes surrounded by several embedded buildings should be worth less than homes near one. Counting embedded buildings within 400 feet of each home (distribution 0/1/2/3+ = [3743, 1605, 1183, 3065]), the estimates are precise zeros:
The assessed-value outcome is itself a model prediction, which can make estimates look more precise than the underlying market signal. As a check, the key specifications are re-estimated on 4,005 qualifying arm's-length sales of universe homes (2017–2025; ln sale price, year dummies; house characteristics are as of the current extract rather than the sale date, so these estimates are noisier by construction). The embedded results are directionally the same and statistically zero, with much wider intervals: -0.1% (95% CI -8.7 to +9.2) within 100 feet in the pooled grid-FE model, and +3.8% (95% CI -1.6 to +9.5) in the embedded ring. The corridor discount is, if anything, larger in transaction prices (-16.8% within 100 feet) than in assessments, consistent with the assessor's model smoothing street-level disamenities.
An independent adversarial audit of the pipeline reproduced the tracked outputs exactly and motivated the checks in this section; all are now part of the pipeline.
Classification variants. The corridor rule was hardened after observing that Austin Blvd's residential frontage misclassified its buildings as embedded, so post-hoc rule choice is a fair concern. Under a commercial-proximity-only rule (no arterial tagging at all) the embedded ring estimate at 0–100 ft is -0.7% (95% CI -1.8 to +0.4); dropping Austin Blvd from the arterial list gives -0.2%. The estimates are not uniformly positive across rules, but they are uniformly small, within about a point of zero.
Spatial inference. Cluster-robust standard errors assume independence across clusters; residual spatial correlation across nearby grid cells could understate uncertainty. Re-inferring the grid-FE split model with Conley spatial-HAC standard errors (Bartlett kernel, 1,000-ft cutoff) changes nothing: embedded 0–100 ft +1.1% (t = 1.23), corridor -4.3% (t = -4.1).
Three limitations bound the interpretation. First, the primary outcome is the assessor's modeled market value rather than transaction prices; the sale-price models of Section 8 point the same direction but with wide intervals, and a fully specified transaction analysis (characteristics as of sale date, repeat-sales structure) remains the natural extension. Second, the design is cross-sectional: it measures whether existing proximity is capitalized, not the effect of new construction, and building locations are not exogenous even within rings; quasi-experimental studies of new multi-family construction and upzoning (Asquith, Mast and Reed 2023; Kuhlmann 2021; Greenaway-McGrevy and Phillips 2023) generally find neutral-to-positive nearby-price effects. Third, the arterial list and thresholds embed judgment; Section 9 shows the embedded estimate stays within about a point of zero under commercial thresholds of 200, 300, 400 feet, a commercial-only rule, and an arterial list without Austin Blvd, though the sign varies across rules.
In the one American suburb's worth of universal parcel data examined here, the claim "multi-family housing lowers nearby single-family values" finds no support in its own strongest form. Across every specification (ring designs around the same building, restricted rings, street-distance controls, alternative classification rules, dose-response counts, spatial standard errors, and recorded sale prices) there is no evidence that homes beside mid-block 2–6-flats, condominium buildings, and townhomes are worth less than observably identical homes two blocks from the same building; the tightest estimates rule out discounts larger than about 0.2% within 100 feet. The one robust proximity discount in the data attaches to arterial streets, is present with no multi-family terms in the model, and absorbs most of the corridor-building discount when controlled directly. For zoning debates, the distinction matters: the housing types that reform would permit on residential blocks are precisely the embedded type, for which no effect on neighbors can be detected.
Asquith, B. J., E. Mast, and D. Reed (2023). "Local Effects of Large New Apartment Buildings in Low-Income Areas." Review of Economics and Statistics 105(2): 359–375.
Greenaway-McGrevy, R., and P. C. B. Phillips (2023). "The Impact of Upzoning on Housing Construction in Auckland." Journal of Urban Economics 136: 103555.
Kuhlmann, D. (2021). "Upzoning and Single-Family Housing Prices: A (Very) Early Analysis of the Minneapolis 2040 Plan." Journal of the American Planning Association 87(3): 383–395.
Linden, L., and J. E. Rockoff (2008). "Estimates of the Impact of Crime Risk on Property Values from Megan's Laws." American Economic Review 98(3): 1103–1127.
Rosen, S. (1974). "Hedonic Prices and Implicit Markets: Product Differentiation in Pure Competition." Journal of Political Economy 82(1): 34–55.
Cook County Assessor. "Assessor – Assessed Values." Cook County Open
Data, dataset uzyt-m557.
https://datacatalog.cookcountyil.gov/d/uzyt-m557
Cook County Assessor. "Assessor – Single and Multi-Family Improvement
Characteristics." Cook County Open Data, dataset x54s-btds.
https://datacatalog.cookcountyil.gov/d/x54s-btds
Cook County Assessor. "Assessor – Parcel Addresses." Cook County Open
Data, dataset 3723-97qp.
https://datacatalog.cookcountyil.gov/d/3723-97qp
Cook County Assessor. "Assessor – Parcel Sales." Cook County Open
Data, dataset wvhk-k5uv.
https://datacatalog.cookcountyil.gov/d/wvhk-k5uv
Cook County GIS. "Cook County Address Points." Cook County Open Data,
dataset 78yw-iddh.
https://datacatalog.cookcountyil.gov/d/78yw-iddh
The pipeline (run.sh, stages s01–s06)
is deterministic: no sampling, no iteration-order dependence; identical inputs
produce identical output. Stage 1 records the source snapshot's size
(6,592,802,816 bytes), modification time, and the SHA-256 of its first 64 MB
in outputs/audit.log, together with before/after counts for every
filter, geocoding match rates by rule, and every regression's coefficients.
All parameters (class lists, thresholds, band edges, arterial list) are in
config.py. This document is generated by s06_paper.py
from outputs/results.json; no quantitative statement is hand-typed.
The source snapshot is rebuildable from the public Socrata datasets listed
under Data Sources.
AI assistance disclosure. The pipeline code, the paper text, and the figures were developed with the assistance of a large language model (Anthropic's Claude), and the code and methodology were independently and adversarially reviewed with a second model (OpenAI's Codex); the review's findings are addressed in Section 9. The data chain and the reproducibility design above exist to keep model assistance out of the numbers: every quantity in this paper originates in the public datasets listed under Data Sources, flows through the deterministic pipeline stages, and is injected into this document programmatically from the pipeline's output file. No number is produced, estimated, or transcribed by a language model, and any reader can rerun the pipeline and reproduce every figure and coefficient exactly.