Vendor evaluation that actually filters
وینڈر کی اصلی جانچ
35 min read
Three ways to see it
A serious vendor evaluation rests on five legs. Technical scoring: does the system work, on our data, in our network, with our latency, in our languages. Financial position: is the vendor solvent for the contract term, can we see audited accounts. References: who has used this product at our scale, what did they really say off the record. Pilots: can we run a paid four to six week pilot on a non-production slice before committing. Inclusion mandates: does the vendor meet the female participation, PWD accessibility, and local capacity requirements we wrote into the RFP.
Way one to score technical capability: insist on a benchmark on your data. Generic benchmarks like MMLU, GLUE, or vendor white papers tell you almost nothing about how a system will behave on a Pakistani complaint corpus mixed with Urdu and Punjabi. Build a 200-sample evaluation set drawn from your own records. Ask each vendor to run their system on it and produce results in a comparable format. Score them on accuracy, latency, and total cost.
Way two to score local capacity and inclusion: write quantitative thresholds, not aspirations. 30 percent female engineers on the delivery team is a threshold. 'We support diversity' is an aspiration. PWD accessibility WCAG 2.2 AA conformance is a threshold. 'Accessible UI' is an aspiration. Pakistani vendors will rise to clear thresholds. Foreign vendors will negotiate around vague ones. Always write the threshold.
Quick check
Quick check: what makes modern AI different from a rule-based program?
The why-tree
Why-tree level one: why score five legs, not just price? Because the cheapest vendor often becomes the most expensive vendor by year three. Total cost of ownership is hidden in capacity, governance, and exit clauses.
Try this with Claude
AI-edge prompt to try: 'Acting as an off-record reference checker, give me ten sharp questions to ask a customer of a major AI vendor that will surface honest signal beyond marketing. Order them from softest to hardest.' Use the list in your next reference call.
Sources
Sources and further reading. Gartner, Magic Quadrant methodology notes (gartner.com). Forrester Wave methodology. World Bank, Pakistan Public Procurement Diagnostic. Transparency International Pakistan, public procurement integrity reports. UNDP Pakistan, women in tech procurement clauses guidance. WCAG 2.2 specifications (w3.org/TR/WCAG22).