Recommended Free Tools
The supported way to retrieve Immowelt listing data is its API, if you are an advertiser with an active Immowelt presentation contract and credentials. That API is for a provider’s own listings—not a general feed for exporting listings from multiple providers. For public pages, first check the live robots.txt, use only accessible paths it does not disallow, and stop if you encounter a login, bot challenge, or other access control. Neither route makes every proposed use lawful: purpose, scale, fields, and publication plans matter.
Choose the route that matches your authorization
| Approach | Who it fits | What to expect | Main constraint |
|---|---|---|---|
| Official Immowelt API | An advertiser with an active Immowelt presentation contract and API credentials | Documented SOAP/XML services for location lookup, filtered listing search, and individual listing details | Not a general multi-provider export feed; the terms restrict third-party retrieval of multiple providers’ listings and pure data export. |
| Public-page observation | A project with a specific, justified need to observe accessible public pages | Page content can change, and HTML parsing is more fragile than a documented service | Check the live robots.txt and applicable terms; do not use disallowed paths, evade controls, or collect contact-form data. |
| Managed extraction provider | A team that has decided external maintenance of selectors and retries is preferable to operating its own parser | A provider may offer an extraction service, but its scope and operating practices need review | A vendor’s description does not authorize your use of Immowelt data. Verify permission, price, service terms, and handling of access controls directly. |
Immowelt’s documented API is the clearest route for an eligible advertiser’s own inventory. Public-page collection is not a substitute for API permission to retrieve other providers’ objects. Before building either route, decide what fields you need, why you need them, who will see the results, and how long you will retain them.
Use the official API for your own listings
Eligibility and terms
AVIV Germany describes the API as an additional service for providers with an active Immowelt presentation contract. API credentials are requested through the provider account. Its terms say third-party retrieval of objects from multiple providers—for example, to display them on a separate marketplace—is not permitted without express consent. The terms also prohibit use for pure data export. Treat these as meaningful limits on the API, not as incidental technical details: access credentials do not by themselves grant permission for every downstream use.
The technical documentation describes a language-independent web service that uses SOAP-capable clients and XML over HTTP. The listed services are LocationService, EstateService, EstateExpose, and CommunicationService. The available documentation summary does not establish an endpoint URL, authentication envelope, operation names, or XML request schema, so do not copy guessed SOAP calls from an unrelated integration. Obtain the current technical documentation and credentials through your provider account, then build against its exact WSDL and schemas.
#1 Best Overall
Documented retrieval sequence
- Confirm access and use. Check that your account is eligible, obtain credentials through the provider account, and verify that your intended storage, rendering, or publication fits the API terms.
- Resolve a location. Use LocationService to turn a town, postcode, or region into a GeoID. Use the returned identifier in a search rather than assuming a location string will be accepted as a filter.
- Search with EstateService. Supply explicit criteria, radius, sorting, and pagination. The documentation sets a maximum of 500 objects per page; use pagination when results exceed the page size.
- Retrieve details selectively. Keep each result’s GUID or Immowelt OnlineID, then request EstateExpose details for a listing when needed. Avoid fetching full exposes when a search result already supplies the fields your authorized use needs.
- Record provenance and refresh. Store the source identifier and retrieval time alongside each record. The documentation warns that an object can be deactivated, so refresh cached records regularly and treat stale status as a real data-quality risk.
- Render within the permitted use. Follow the API terms on attribution and publication when displaying your own inventory. Do not assume the API grants permission to republish listings on another property portal.
What to store
Keep a minimal record that supports your purpose and lets you refresh or remove a listing accurately. A practical starting point is the Immowelt identifier, requested fields, retrieval timestamp, and the listing’s current state as returned by the service. Add price, floor area, location, or expose details only where they are needed and permitted. Treat a missing result, deactivated object, and failed request as different states rather than silently preserving an old listing as current.
Public-page collection: a cautious fallback
The live Immowelt robots.txt disallows internal endpoints, maps, booking and contact paths, previews, parameterized classified-search and classified-map URLs, and classifiedList. It also disallows specified tracking and backend paths, blocks twiceler and NerdByNature.Bot, and sets an AhrefsBot crawl-delay of 50. It publishes a sitemap index. These rules can change; retrieve and review the current file before each crawl run. A path being absent from a disallow rule is not a legal permission or a guarantee that automated collection is allowed.
Rank #2
Safe planning checklist
- Build an allowlist from public pages that are accessible and not disallowed under the current robots.txt. Do not crawl the disallowed search, map, contact, booking, preview, or backend paths.
- Review Immowelt’s current terms and your project’s legal basis separately from robots.txt. Robots rules guide crawler behavior; they are not a license to collect or reuse content.
- Use measured request rates, identify your crawler honestly, cache conservatively, and stop if the site presents a login, CAPTCHA, bot challenge, or another access control. Do not rotate identities or otherwise work around blocks.
- Do not submit contact forms, invoke communication endpoints, or collect names and contact details unless essential and lawful for the project.
- Re-check page structure and listing status over time. HTML selectors can break after a site change, and listings can be removed or deactivated.
A conservative Python starter
This runnable example checks robots.txt for a URL supplied by you before requesting that one page. It extracts only JSON-LD objects already present in the returned HTML and prints them as JSON; it does not claim Immowelt pages expose listing data in that format. It is a starting point for a page you are authorized to inspect, not a way around blocked paths or a ready-made Immowelt parser. Install dependencies with python -m pip install requests beautifulsoup4.
import json
import sys
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
USER_AGENT = "AuthorizedListingResearchBot/1.0 (contact: webmaster@example.org)"
TIMEOUT_SECONDS = 20
def robots_allows(url):
parts = urlparse(url)
if parts.scheme not in ("http", "https") or not parts.netloc:
raise ValueError("Provide a complete http or https URL")
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
response = requests.get(
robots_url,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
response.raise_for_status()
parser = RobotFileParser()
parser.set_url(robots_url)
parser.parse(response.text.splitlines())
return parser.can_fetch(USER_AGENT, url)
def main():
if len(sys.argv) != 2:
raise SystemExit("Usage: python inspect_page.py https://example.org/public-page")
url = sys.argv[1]
if not robots_allows(url):
raise SystemExit("Robots rules disallow this URL for this user agent")
response = requests.get(
url,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "text/html" not in content_type.lower():
raise SystemExit(f"Expected HTML, received {content_type or 'unknown content type'}")
soup = BeautifulSoup(response.text, "html.parser")
results = []
for tag in soup.find_all("script", type="application/ld+json"):
try:
results.append(json.loads(tag.string or tag.get_text()))
except json.JSONDecodeError:
continue
print(json.dumps({"url": response.url, "json_ld": results}, ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
Replace the example user-agent contact with a real monitored address before any use. The standard-library robots parser is a useful preliminary check, not a substitute for reviewing the live rules: implementations can differ in how they handle edge cases. Review the exact target path and any relevant terms yourself. If robots.txt cannot be retrieved, the example stops on the HTTP error instead of proceeding as if the URL were allowed.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #3
Make the data handling proportionate
AVIV Germany identifies itself as the controller for immowelt.de. Its privacy notice says the site processes IP address, URL, date and time, browser version, operating system, cookies, and usage information for operation, analytics, and IT security and bot protection. Your own collection creates separate questions about the data you gather, the purpose, retention, and sharing.
- Maintain a data inventory: list the fields collected, their source, purpose, recipients, and retention period.
- Minimize personal data. Listing price, floor area, and location may be relevant to an analysis; names and contact details often are not.
- Set a retention and refresh policy. Remove records that are no longer needed, and avoid presenting stale listings as active.
- For commercial or large-scale projects, document the purpose and legal basis with qualified counsel. General guidance cannot decide whether a specific collection plan is lawful under German or EU law.
Reliability, freshness, and cost trade-offs
The API offers a documented SOAP/XML contract, identifiers, search filters, sorting, and pagination, but requires eligible credentials and remains subject to the API terms. Public HTML may be accessible without credentials, but its structure can change and the robots rules restrict important paths. A managed extraction service can shift selector maintenance and retries to a vendor, but it does not resolve your authorization, privacy, or republication obligations.
For a maintainable pipeline, store source identifiers and retrieval timestamps, use conservative request pacing, retry only transient failures with backoff, and distinguish transport errors from a listing that no longer appears. Do not retry a bot challenge as though it were a transient network fault. For cached data, set a refresh interval appropriate to the use and remove or revalidate inactive objects. The API documentation specifically warns that an object may be deactivated; no universal refresh interval is established here.
Troubleshooting common problems
| Symptom | Likely cause | Safe next step |
|---|---|---|
| Provider account cannot obtain credentials | The API is limited to providers with an active Immowelt presentation contract, or the account setup is incomplete. | Confirm eligibility with the provider account. Do not search for another provider’s key or reverse-engineer a private endpoint. |
| Location search returns no GeoID | The input may be ambiguous, misspelled, or not the expected location form. | Use LocationService as documented, try a precise postcode or town name, and inspect the service response before continuing. |
| Search results stop before all matches | Pagination may not be advancing, or the per-page maximum has been reached. | Follow the documentation’s pagination fields and request subsequent pages. The documented maximum is 500 objects per page. |
| An identifier no longer returns an expose | The listing may have been deactivated or changed since the earlier search. | Refresh through EstateService and update the record’s status instead of treating an old cached expose as current. |
| A public-page request is blocked or challenged | The path may be disallowed or the site may be applying access controls. | Stop. Re-check robots.txt and terms; do not evade the block, automate contact functions, or keep retrying a challenge. |
| The parser returns no listing fields | The page may not publish structured data in the format the example checks, or its HTML may have changed. | Inspect only an authorized accessible page, verify its content type and current structure, and avoid assuming selectors or fields that the page does not expose. |
| CSV output appears incomplete or stale | Pagination, field selection, or refresh handling may be incomplete; listings may also have been deactivated. | Trace records by source identifier and retrieval time, verify each page of results, and refresh status before treating the export as current. Confirm that the intended export use is permitted. |
Or skip the browser setup
ScreenshotNeo can capture a permitted public page as an image or PDF, but a screenshot is not structured listing data and does not replace the Immowelt API. For a visual record of a page you are allowed to access, one request can return a capture:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.immowelt.de -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These capture features do not authorize scraping or republication of Immowelt data.
Sign up for ScreenshotNeo free to get 1,000 screenshots a month with no card.
Further reading
Ryan Mitchell’s Web Scraping with Python, 2nd Edition (April 2018) covers BeautifulSoup, crawler design, Scrapy, and data storage. It can help with general parsing and pipeline concepts, but it does not supply Immowelt-specific authorization or current page selectors.
Frequently Asked Questions
Does Immowelt provide an API for any developer to export listings?
No general multi-provider export feed is established by the API terms. The documented route is for eligible advertisers’ own inventory.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Can a screenshot be converted into a reliable CSV of listings?
A screenshot is a visual capture, not a structured listing response. Use an authorized data source for fields such as price, floor area, and identifiers.
Does an allowed robots.txt path mean I have permission to reuse its listings?
No. Robots rules guide crawler access; they do not establish legal permission to collect, retain, or republish content.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

