After two quarters of investment into integrating Fundamental data in Q, specifically with the goals of making Q a serious targeting & benchmarking tool, we learned a great deal about how to improve Q’s quality of responses.
Users have to often explain context, re-prompt with guidance and re-frame their goals — otherwise the output lags behind what any IR professionals can produce. This makes Q feel more like a search interface than a strategic partner.
This leads to inconsistent outputs, limited repeatability, and significant under-leveraging of the proprietary IR datasets Q4 provides access to: ownership data, financial data, peers, and engagement history.
Users have to often explain context, re-prompt with guidance and re-frame their goals — otherwise the output lags behind what any IR professionals can produce. This makes Q feel more like a search interface than a strategic partner.
This leads to inconsistent outputs, limited repeatability, and significant under-leveraging of the proprietary IR datasets Q4 provides access to: ownership data, financial data, peers, and engagement history.
Desired Outcomes & Measurement
IR professionals can execute complex, high-value tasks through pre-built Agents, Workflows & Skills — without needing to prompt engineer
IROs today spend significant time constructing the right prompt before they can get a useful output. A technically skilled user gets a better result than a novice user — not because Q is smarter for one versus the other, but because prompt quality drives output quality. This is a product failure, we cannot expect our users to be prompt-engineers.
Skills should simplify that complexity. The user selects or triggers a skill, provides the minimum necessary context, and receives a structured, immediately actionable output — grounded in Q4's proprietary data & industry expertise — on the first attempt.
Success metric: 50%+ reduction in the number of prompt turns required before a user accepts or acts on an output, for skill-supported use cases.
IROs today spend significant time constructing the right prompt before they can get a useful output. A technically skilled user gets a better result than a novice user — not because Q is smarter for one versus the other, but because prompt quality drives output quality. This is a product failure, we cannot expect our users to be prompt-engineers.
Skills should simplify that complexity. The user selects or triggers a skill, provides the minimum necessary context, and receives a structured, immediately actionable output — grounded in Q4's proprietary data & industry expertise — on the first attempt.
Success metric: 50%+ reduction in the number of prompt turns required before a user accepts or acts on an output, for skill-supported use cases.
Output quality is consistent, explainable, and trustworthy
Every skill deployment should produce outputs that meet a defined quality bar out of the gate — not make “users as our testers”.
This requires an eval framework that defines what "good" looks like for each skill, a golden set of expected outputs to test against, and a gate that prevents deployment of skills that don't meet the standard. The system should make it structurally impossible to publish a skill that hasn't passed the eval gate — similar to a CI/CD framework.
Success metric: Eval pass rate defined per skill before launch; post-launch output acceptance rate tracked via PostHog; no skill deploys without a passing eval threshold.
Every skill deployment should produce outputs that meet a defined quality bar out of the gate — not make “users as our testers”.
This requires an eval framework that defines what "good" looks like for each skill, a golden set of expected outputs to test against, and a gate that prevents deployment of skills that don't meet the standard. The system should make it structurally impossible to publish a skill that hasn't passed the eval gate — similar to a CI/CD framework.
Success metric: Eval pass rate defined per skill before launch; post-launch output acceptance rate tracked via PostHog; no skill deploys without a passing eval threshold.
Agent, Skills & Workflow capability can be extended without requiring an engineering sprint for every change
The business logic inside skills needs to be close to the source. IR SMEs understand targeting nuance, investor segmentation logic, and tone requirements far better than R&D does. A framework that requires full engineering involvement for every skill update is too slow and too insulated to keep pace with customer needs.
The Skills Toolkit should enable a defined set of authors — Q4 IR SMEs, and eventually advanced customer-facing teams — to create, test, and publish skills within a governed workflow. R&D's role is to "owns the toolkit framework that skills are built on", while others take on the role of “creating the right skills”.
Success metric: Time from skill idea to live deployment (end-to-end cycle time); proportion of skill updates that don't require any engineering input.
The business logic inside skills needs to be close to the source. IR SMEs understand targeting nuance, investor segmentation logic, and tone requirements far better than R&D does. A framework that requires full engineering involvement for every skill update is too slow and too insulated to keep pace with customer needs.
The Skills Toolkit should enable a defined set of authors — Q4 IR SMEs, and eventually advanced customer-facing teams — to create, test, and publish skills within a governed workflow. R&D's role is to "owns the toolkit framework that skills are built on", while others take on the role of “creating the right skills”.
Success metric: Time from skill idea to live deployment (end-to-end cycle time); proportion of skill updates that don't require any engineering input.