License real short-term-rental data.
Built for training AI.
We run a live Atlanta short-term-rental portfolio and the software that operates it — and we capture rich, labeled data most teams can't get anywhere else: cleaning and turnover photos, real guest conversations, maintenance records, revenue and occupancy data, and document-extraction corpora. If you're training a model and this is useful, let's talk.
What we capture
Datasets from a real operation
Every dataset below comes from running actual short-term rentals — not a scraped or synthetic sample. Availability and detail are scoped per engagement.
Cleaning & turnover capture
AI-generated per-property checklists with photo documentation of every room on every clean, plus a full cleaner audit log of each job.
Useful for: Vision & robotics for cleaning, task verification, before/after models
Move-out condition photos
Guided move-in vs. move-out photo inspections with paired before/after images and structured condition evidence.
Useful for: Damage detection, condition assessment, paired-image vision
Airbnb guest-messaging threads
Real host↔guest conversations across inquiries, bookings, and stays — a category with no public API, so it's rarely available.
Useful for: Conversational & hospitality-domain LLM training
Maintenance records + photos
Maintenance tickets with structured scope of work, vendor and quote data, issue category, and attached photos.
Useful for: Repair triage, structured extraction, work-order models
Revenue & occupancy data
Monthly per-property performance — ADR, occupancy, nights booked, net figures — alongside competitor benchmarks.
Useful for: STR pricing, demand & market-forecasting models
Document-extraction corpus
Raw scraped public records paired with structured, extracted JSON fields and a scored, ruled label — the messy-doc-to-clean-fields pairing extraction teams need.
Useful for: OCR & document-extraction model training
Listing-quality labels
Structured listing attributes paired with a deterministic quality score across photos, reviews, pricing, and copy.
Useful for: Quality-scoring & listing-ranking models
Who it's for
If you're training on the physical world of hosting
Foundation & vertical models
Ground a model in how real short-term rentals are operated, day to day, across a live portfolio.
Vision & robotics
Room-level cleaning and condition imagery with paired before/after and task labels.
Document extraction
Messy public records paired with the clean, structured JSON they were turned into.
Conversational AI
Real hospitality conversations for guest-service and agent training.
How it works
From inquiry to delivery
Tell us what you need
Send an inquiry with the datasets, task, and rough volume you're after.
Scoping call
We confirm what we can deliver, in what format, and talk through fit and pricing.
Data-use agreement
We sign a data agreement covering scope, use, and privacy before anything moves.
De-identified delivery
You receive a de-identified dataset in the agreed format; record-level access stays under the agreement.
De-identified by default
Data is shared de-identified and aggregated by default. Record-level or raw access is available only case by case and only under a data-use agreement that governs scope, use, and privacy. We take the privacy of our guests, owners, and vendors seriously — nothing moves without the right terms in place.
Request access
Start a data conversation
Tell us what you're building and which datasets look useful. We'll follow up to scope a partnership.
FAQ
Data licensing questions
What short-term-rental data can I license?
Real operational data from a live Atlanta short-term-rental portfolio and its software stack: cleaning and turnover capture (per-room checklists and photos), move-out condition photos, Airbnb guest-messaging threads, maintenance records with photos, monthly revenue and occupancy data with competitor benchmarks, a public-record document-extraction corpus, and listing-quality labels.
Is the data de-identified?
Yes. Data is shared de-identified and aggregated by default. Record-level or raw access is available only case by case and only under a data-use agreement that governs scope, use, and privacy. Some datasets contain personal information at the source and are never shared without that agreement.
What formats do you deliver in?
We scope format with you on the call — typically structured exports (CSV/JSON/JSONL) for records and conversations, and image sets with accompanying labels for the vision datasets. Tell us what your pipeline expects and we'll work to it.
How is licensing priced?
Every engagement is bespoke — pricing depends on which datasets, the volume, the format, and how the data will be used. Send an inquiry and we'll quote it on the scoping call. We don't publish fixed prices.
Can we get a sample before committing?
In most cases, yes — a small de-identified sample can be arranged under the data-use agreement so your team can evaluate fit before a full license. Mention it in your inquiry.
Who is this for?
AI labs, research groups, and data platforms training foundation or vertical models, computer-vision and robotics teams working on cleaning or condition assessment, document-extraction and OCR teams, and anyone building conversational or hospitality-domain AI.