Description
Solveion is an independent AI consultancy. We build custom AI systems for clients: assistants, automations, and retrieval systems. Those systems need training and test data before real client data exists, or where real data cannot leave the client. I need someone to generate that data. This is a developer role, not a labeling role. The work varies by project, but it looks like this: - Generate synthetic records to a schema I give you: support tickets, invoices, contracts, intake forms, chat transcripts - Build the generation pipeline in Python, with a seed and a config so the same run reproduces - Write matching evaluation sets: the input, the correct answer, and the reason it is correct - Build in the messiness on purpose. Typos, missing fields, contradictions, and edge cases, because clean data teaches a system nothing about production. - Document what you generated, how, and what the distribution looks like Volume and cadence change project to project. Right now I expect around 40 hours a month across two or three projects. What I need from you: - You have generated synthetic data before, whether with LLMs, Faker, SDV, or your own scripts, and you can explain the trade-offs between those - You write Python I can read and hand to someone else - You understand why a synthetic set that is too clean or too uniform is worse than no set at all - You will sign an NDA. Schemas and business rules come from client work. Useful but not required: experience with retrieval and RAG eval