A hands-on role for someone who writes with precision and can tell a great customer conversation from a mediocre one. You'll shape how our AI talks to real people — no computer science degree required.
We're a small team building two products. Skillify runs career workshops and mentorship programs for students — we've reached thousands of students across 220+ schools. Kinship Labs (kinshipcomms.ai) is the AI-powered relationship management platform that powers Skillify and helps other organizations maintain personalized, human connections at scale across SMS and email. The conversations you'll work on power both.
Every day, our AI agents hold real conversations with real customers — answering questions, following up, scheduling, and keeping relationships warm. Your job is to make those conversations excellent. You'll write and refine the prompts behind the agents, turn real conversations into test scenarios and evals using our built-in evaluation platform, and review live conversations to judge tone, accuracy, and empathy.
This is a judgment-and-writing role, not a coding role. Our platform lets you create evals without writing a line of code: take any real conversation, freeze it into a test scenario, define what a good response looks like in plain English, and run the whole suite against a new prompt to see what passes and what breaks.
Work environment: full-time office setting with a small, collaborative team. You'll have a dedicated workstation, access to professional AI tooling (Claude, OpenAI, Grok, etc.), and direct mentorship from the team building the platform.
What you'll actually do
Prompt engineering & optimization (40%)
- Write and iterate on the production prompts behind our customer-facing AI agents
- Tune agent voice and tone so every message sounds personal, warm, and on-brand
- Sharpen how agents handle the hard moments: complaints, confusion, sensitive topics, and knowing when to hand off to a human
- Compare prompt versions head-to-head and promote the winners to production
- Document what works — build a library of effective prompt patterns for the team
Building & running evals (35%)
- Turn real customer conversations into frozen test scenarios using our in-platform eval tools
- Write eval criteria in plain English — "Does the response acknowledge the customer's concern before recommending anything?"
- Run eval suites against prompt changes and read the results: what passed, what broke, and why
- Grow the eval library over time so every failure we've ever seen stays fixed
Conversation review & quality (25%)
- Review live AI customer conversations across SMS and email for tone, accuracy, and empathy
- Spot failure modes and edge cases — the awkward reply, the missed cue, the moment that needed a human
- Turn what you find into new test scenarios, new evals, and better prompts
- Report conversation quality trends back to the product and engineering teams
What you'll learn
By the end of this internship, you'll be able to:
- Engineer production prompts — write, test, and iterate on prompts that perform reliably across thousands of real conversations
- Evaluate AI systems rigorously — design eval scenarios and criteria that measure what actually matters, one of the most in-demand skills in AI right now
- Judge conversation quality like a professional — develop a trained eye for tone, empathy, escalation, and brand voice at scale
- Work with frontier AI models — hands-on daily experience with Claude, GPT-5, and Grok in a real production setting
- Operate like a product team — track work in Linear, communicate findings clearly, and turn observations into shipped improvements
What we're looking for
Required
- Currently enrolled in or recently completed a Bachelor's or Master's program — any major. Communications, English, psychology, linguistics, business, and technical fields are all welcome
- Exceptional written communication — you notice when a sentence lands wrong, and you can explain why
- Strong judgment about conversations: tone, empathy, and when a situation needs a human
- Genuine curiosity about AI — you already use tools like ChatGPT or Claude and think about why they respond the way they do
- Detail-oriented and systematic — you can review dozens of conversations without your standards slipping
Nice to have
- Customer service, support, hospitality, or teaching experience — anywhere you've handled real people in real conversations
- Experience experimenting with prompts, custom GPTs, or AI writing tools
- A writing, editing, or content background
- Comfort with spreadsheets and light data analysis
Tools you'll work with
- AI models: Claude (Anthropic), GPT-5 (OpenAI), Grok (xAI)
- Eval platform: Kinship's built-in test scenario & evaluation tooling — no code needed
- Project management: Linear
- Communication: Slack
Mentorship & path to full-time
You'll get daily check-ins with progress review, weekly 1:1 mentorship sessions, and direct access to the founders and engineers building the platform — your findings will shape the product in real time.
This internship is designed with a potential transition to a full-time role in mind. At the end of the 3-month program, strong performers may be offered a full-time position as a Prompt Engineer or AI Quality Analyst. Final hiring depends on mutual fit, business needs, and available funding at the time of evaluation.
Sound like you?
If you read this and thought "I was literally born for this" — we want to hear from you.