Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

The Zillow Prize was a completed, two-phase Kaggle competition—not an open contest you can still enter. In the qualifying round, participants predicted Zillow Zestimate errors for Fall 2017 home sales in Los Angeles, Orange and Ventura counties. An invitation-only final phase changed the target to actual sale prices and compared submissions with a Zillow-built benchmark. Kaggle reported that Team ChaNJestimate won with a 0.12110 score versus Zillow’s 0.14084 benchmark, describing the result as more than 13% better.

What the Zillow Prize asked competitors to predict

The public competition centered on logerror = log(Zestimate) - log(SalePrice). A positive value meant Zillow’s Zestimate was above the eventual sale price; a negative value meant it was below. Competitors used property characteristics and transaction information to predict that residual for Fall 2017 sales.

The supplied property list covered three California counties: Los Angeles, Orange and Ventura. Training materials included 2016 property data and transaction information, while evaluation depended on subsequent sales. The objective was therefore not simply to estimate a home’s value from scratch; it was to identify where and when Zestimate errors tended to occur.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The contest was launched by Zillow and Kaggle on May 24, 2017. Kaggle’s displayed qualifying-round close date was January 10, 2018. Zillow’s contemporaneous announcement described the schedule somewhat differently, so the Kaggle competition page is the clearest reference for the displayed deadline.

How the two competition phases differed

Aspect Qualifying round Final phase
Eligibility Public Kaggle participation subject to the competition rules Restricted to qualifying top performers; Zillow could select up to the top 100 submissions for possible advancement
Primary target Zestimate log-error: log(Zestimate) - log(SalePrice) Actual sale price prediction
Data and modeling Provided assessor, property and transaction data New data sources and engineered features were encouraged
Evaluation Qualification against later Fall 2017 sales Later sales evaluation against a competition-specific Zillow benchmark
Operational obligations Compliant submissions and team registration Additional participation terms, limits on sharing outside teams, and delivery of final model software and documentation for a prize-winning solution

The final-phase benchmark was not simply the Zestimate shown on Zillow’s website. Zillow said it was a modified Zestimate trained on the same final-round data. That distinction matters: beating the published benchmark meant outperforming the contest’s defined baseline under its own data and evaluation setup, not proving that every Zestimate was inaccurate in every market.

A practical way to approach the historical qualifying task

  1. Interpret the sign correctly. Keep the target definition visible in your notebook. A positive log-error indicates overestimation; a negative value indicates underestimation.
  2. Audit the supplied data. Inspect property records, transaction dates, missing values, duplicated parcels and county-specific fields before fitting a model. Treat the three counties as related but not interchangeable markets.
  3. Respect time. Build validation around the competition’s sales timeline rather than randomly mixing future transactions into training folds. The score was determined on subsequent sales, so a time-aware split better reflects the information available at prediction time.
  4. Engineer defensible features. Use property attributes and local-market signals available in the permitted data. Features should be calculated without using information that would only become known after the prediction date.
  5. Check the submission contract. Confirm team membership, file format, prediction identifiers and deadlines against the rules. A strong model that violates an eligibility or submission condition cannot qualify.

These steps describe a sound interpretation of the published setup; the official material reviewed for this guide does not establish the winning team’s complete feature list, validation design or ensemble recipe.

What changed when finalists predicted sale prices

The second phase was a different modeling problem, not merely a larger version of the first leaderboard. Zillow described it as predicting actual sale prices while allowing innovative data sources and feature engineering. Finalists therefore had to reconsider the target, the sources they could lawfully use, and how to compare their predictions with Zillow’s contest benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The rules also introduced delivery and governance requirements. Only qualifying entries could advance to possible participation, and a second-round prize winner had to provide working model software and documentation. Team and sharing restrictions applied, so collaboration had to remain within the permitted structure.

The reported winning result

Kaggle’s winner announcement named Team ChaNJestimate as the winner and reported:

  • 0.12110: the team’s reported final score.
  • 0.14084: Zillow’s benchmark score cited in the same announcement.
  • “Over 13% better”: Kaggle’s characterization of the improvement over that benchmark.

Those are historical competition figures as reported by Kaggle. The announcement establishes the result and comparison, but it does not, by itself, document a reproducible winning pipeline. Claims about a specific ensemble, feature inventory or validation trick require the team’s own technical write-up or code.

Why Zillow sponsored the challenge

Zillow’s 2017 materials framed the contest as a search for hyperlocal data and new algorithms. Stan Humphries, Zillow Group chief analytics officer and creator of the Zestimate, wrote: “We’re particularly excited about the exploration of more hyperlocal data and algorithms, a task well-suited to highly distributed, crowd-sourced efforts.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In another contemporaneous announcement, Zillow attributed to Humphries this description of the next stage of innovation: “While that error rate is incredibly low, we know the next round of innovation will come from imaginative solutions involving everything from deep learning to hyperlocal data sets — the type of work perfect for crowdsourcing within a competitive environment.” The wording reflects Zillow’s motivation at the time, not a current accuracy claim.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Historical Zestimate figures and their limits

Zillow’s 2017 announcements supplied several context-specific statistics:

  • Zillow said it published Zestimates for more than 110 million homes.
  • The company said Zestimate calculations used 7.5 million statistical and machine-learning models.
  • Zillow reported a 5% U.S. median absolute percentage error, improved from 14% in 2006.
  • In a study of 2016 transactions listed for sale on Zillow, Humphries reported 3.5% Zestimate error versus 2.5% error for the listing price.

These figures describe the populations and definitions stated by Zillow in 2017. The 3.5% study covered listed-for-sale transactions and was described as more accurate than the overall set; it should not be generalized to every home or treated as a current national rate. The differing announcements also used different populations and error descriptions, so they should not be merged into one universal accuracy number.

Can you enter the Zillow Prize now?

No. Kaggle’s overview lists the qualifying round as closing in January 2018, and the rules describe a second phase that ended in 2019. The useful takeaway today is the competition design: a public residual-prediction qualification followed by a restricted sale-price challenge with a sponsor-defined benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a modern reader can learn from the contest

  • Define the target precisely: predicting a residual around an existing model is different from predicting the underlying value.
  • Match validation to deployment: later-sale evaluation makes temporal leakage especially dangerous.
  • Separate leaderboard success from reproducibility: a reported score does not reveal the complete winning method.
  • Read the rules as part of the technical problem: eligibility, data-sharing limits and model-delivery obligations can determine whether a submission is prize-eligible.
  • Interpret benchmarks narrowly: beating a contest-specific Zillow model is evidence about that benchmark and evaluation period, not a blanket judgment about all Zestimates.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.