InsurTalk: Implementation of AI Augmented Underwriting for Specialty Insurance with Paul Butler, CTO at Hiscox London Market
As artificial intelligence and machine learning continue to evolve, their applications within the insurance sector are expanding rapidly, offering new possibilities for enhancing efficiency, decision-making, and customer service.
In this episode we have a pleasure to talk to Paul Butler, Chief Technology Officer at Hiscox London Market. Our conversation centers on the innovative work his team has done in implementing AI within the underwriting process. He shares insights on the challenges and successes experienced along the way, offering a detailed look at how these technologies are being integrated into the London Market, the impact they have on speed and accuracy, and the future of AI in the insurance industry.
Piotr Piękoś: Welcome to another episode of IT Insights InsurTalk by Future Processing. My name is Piotr Pinos [transcribed as p penos in source], and today I’ll be your Piotr Piękoś of the show. I have a great privilege to interview Paul Butler, Hiscox London Market Chief Technology Officer. Paul and his team had made remarkable advances in applying machine learning recently. The team had successfully implemented large language models to carrier’s arguably most critical function: underwriting. Today we are going to talk about the practical implementations of AI augmented underwriting for a major Lloyd’s of London carrier, the lessons learned that come from those implementations, and the blueprint of the technology organization that can accommodate the benefits yielded by artificial intelligence. Welcome to the show, Paul.
Paul Butler: Thank you. Thank you for agreeing to do this.
Piotr Piękoś: We’re shooting today in Cornelia Parker Room at Hiscox headquarters, and to me, it’s been always very interesting how Hiscox’s culture blends art, engineering, and insurance together. And I know that for you, the art is also something close to your heart. So maybe you could tell our audience what role music plays in your life.,
Paul Butler: Yeah, sure. Yeah, I mean, I’m a massive fan of music, all types of genre. I’m not particularly tied to anything in particular. I’ve played instruments most of my life. From a young child, I used to play the clarinet, and more recently, I’ve sort of got a slight addiction to buying vintage synthesizers and playing lots of synthesizers. But yeah, music’s really important. It’s a way for me to decompress. You know, work is very fast-paced, can be stressful at times, but getting into music, it can just take you away from all of that, especially if you’re creating it and making it.
Piotr Piękoś: In what ways do you think it impacts your decision making?
Paul Butler: I think from doing music as a child, I think it helped me to sort of get into patterns and logical thinking because originally I was an engineer. Most of my career I’ve been sort of an architect, an engineer, getting involved writing code. So I think it sort of tuned my brain to be sort of quite logical in terms of thinking, in sort of ways of working. Music’s… it’s quite a sort of a good analogy. So, you know, if you’re doing classical music with a big orchestra, it sort of feels like how you might approach running technology. You know, you’ve got a playbook, which is what the composer’s written down—so the documentation—you understand what skills you need in that orchestra to sort of perform that piece of music. So when you’re sort of running technology, you have knowledge sharing, you have documentation, you have different people with different skills, and you bring them all together. So [that’s] a good analogy. And then, if you’re being sort of innovative and creative, then I think of jazz. You know, there’s no sort of hard-written manuscripts and you have these smaller agile teams of skilled people that have the freedom to experiment and innovate. So music sort of touches on the two sort of things we do at Hiscox. You know, partially our team is running the stuff we have already, the products and the services, and the other part about what we do is innovate and do new creative things.,,
Piotr Piękoś: Certainly you’ve added a new instrument to your technology orchestra recently. So could you please give our audience a short summary of the AI work you and your team have done at London Market, particularly in the underwriting segment?
Paul Butler: I mean, we we’ve been working on AI for some time [at] the London Market, probably about four years. That’s when we sort of created a brand new data science team. We’ve got a fantastic head of data science, Matthew Hudson, who’s built a really good chapter in our in our team. And we spent the first sort of couple of years building sort of neural networks, using a lot of back data to train models to effectively take submission emails and the attachments that you tend to get in submission emails, like schedule of values, and automatically strip that data out, transform it, geocode it, and sort of make it available for the underwriters to to start using.
Before we used neural networks to do that, we had third-party operational people in India, about 20 people, and it would take them about two or three days to cleanse that data, to transform and geocode that data, and get it back to us. So the models are doing this in seconds. So we’ve reduced that part of the value chain down from days to seconds. But also those models are more accurate as well. So they they sort of got accuracy rates of around 98-99%, whereas what we were getting back from human beings was about 95% accurate, which is still pretty good, but you know, we’ve improved that as well as the speed.
More recently, we’ve been working with Google Cloud. We started sort of talking to them about this time last year. We got quite excited about, you know, could we do something around lead algorithmic underwriting. And we did a proof of concept with them towards the end of last year where we were using their large language models. So what we were doing was we were taking some of what we’d already spent two years building—that, what we call Halo—and then sort of seeing if we could also make use of some large language models in certain parts of that value chain to make further improvements in terms of speed and performance. And we’ve had some great results in in using those as well.,
Piotr Piękoś: It seems to me that you went from proof of concept to actual implementation fairly quickly when it comes to the insurance world. So how do you think the work done by your team today will affect the underwriting in the year to come?
Paul Butler: That’s a good question. I think, you know, we’re calling it augmented underwriting for a reason. So we’re not looking to replace the underwriters; they’re still very key to what we do. We’re in specialist big insurance, so I don’t think we’re ever probably going to get to a point where we’re going to trust large language models, and the fact that they can have bias and hallucinate, to make decisions. So we’re sort of really assessing the risks of using these models and we’re only really going to be prepared to use them where we see the risks quite minimal and the results and the benefits are quite high. So, for us, it’s about getting the information that the underwriters need to make a good decision in front of them as quickly as possible.,
And today, you know, there’s lots of manual processes, lots of people involved to get all the information that we’re getting over from the brokers, doing all the sort of modeling in terms of, you know, does that fit our appetite, is that within the the sort of exposure boundaries that we have set out, and then, you know, how are we going [to] price that risk? All of that stuff takes time. So what we’re trying to do is minimize that time, put all that information that we’ve extracted and packaged together in front of the underwriter for them to make the decision. And what we found with the work we’ve done with Google is we’ve sort of got a three-day process down to about five minutes. And most of that process was accumulating all that information, doing the modeling, coming up with a price, understanding our exposure. And we’ve reduced that bit, but we [are] still now putting all those facts and information in front of the underwriter to make the final decision.,
Piotr Piękoś: Okay, so most of the effort is focused on productivity improvements and assisting the decision making. If we were to look let’s say far into the future, what do you think would need to happen, how the LLMs would need to evolve in order to give them the power to make decisions, if at all?
Paul Butler: Yeah, I I still don’t think you’d be trusting a sort of GenAI model to make those sort of decisions. It’d be far too risky for us and our customers. So again, we’ve sort of we’ve spent a lot of time—I mean, we haven’t gone into production yet on this. We’re working with a particular line of business, sabotage and terrorism, and we’re hoping to go into production soon. The data science piece has been done, but what we’ve spent a lot of time focusing on more recently is the governance frameworks, the controls we’re going to need to have in place. How are we going to sort of monitor these models, make sure that they’re not drifting, make sure that they are accurate, they’re not hallucinating or coming up with the wrong answers. So we’re investing a lot more time in that before we’re prepared to go into production.,
So sort of getting back to your question, I don’t think those types of models are ever going to sort of replace the the human decision making. I can’t see that at the moment. I’m not saying that there might be some future models that are more sort of tuned and specifically created for certain use cases that might be something that we could make use of. But at the moment we’re using things like Gemini, and we’ve tried using things like ChatGPT in the past; they’re just too generic to to replace highly skilled people. But yes, you know, in the future there might be an opportunity to, at least for certain lines that are less complex where we can sort of perhaps expand what we do in terms of appetite, where we might go for more volume and lower value risks, because we could maybe automate that—maybe at that point we can use some form of model that helps do that.,
Piotr Piękoś: When it comes to the specialty lines of business like sabotage that you mentioned, currently the key benefits are related to productivity improvements. What other competitive advantages can generative AI give to the Lloyd’s capacity provider?
Paul Butler: Yeah, so in that particular line, the the actual real value to the business line is the the speed to market. So yes, getting it… getting the turnaround from of submission coming in from the broker, getting the quote back to the broker from three days to five minutes, which is where we’re at. So when we get the quote back to the broker quickly, um, we tend to win the business at the price that we feel is the right price.,
Piotr Piękoś: Okay, so it certainly improves the quality of service that brokers receive from from a carrier. Do you think there are any additional advantages that brokers can benefit from when interacting with a carrier that employed AI augmented underwriting?
Paul Butler: Yeah, I mean I think the benefit for them is that they they get a quote back quickly. You know, and that’s going to benefit the customer as well. You know, I think there’s also that sort of a other use case of auto-follow that we’ve seen with Ki, and some brokers are sort of coming up with their own technology around that as well. I think there’s other benefits and other use cases in terms of what we can do with AI, not always large language models but traditional machine learning as well. In particular, you know, thinking about claims, we could maybe use satellite imagery to quickly assess if a damage to a property is genuine, you know, see the before and after shots and get the machine to make the decision to just pay the claim quickly. So obviously the benefit there is for the customer, right? Absolutely, where we just pay out quickly, you know, there’s no messing, it’s a fast decision-making thing.,
Piotr Piękoś: It’s really interesting how you mentioned the ESG, because in my mind it already sort of comes into, um, in one bit with the energy consumption of the AI modern AI models, right? So I wonder if it’s going to be a net positive or not at the end.
Paul Butler: Yeah, I guess it, you know, depends on different vendors and how they run their data centers. But Google are very good in that regard. You know, a lot of their data centers are carbon neutral, running on renewable power.
Piotr Piękoś: It’s quite surprising I think that those factors are are playing a substantial role in implementing the technology. Are there any other surprises that you’ve encountered when going through the journey of of AI implementation?,
Paul Butler: Yes, yeah, plenty of surprises. Um, yeah, these large language models are gnarly things. They, um, they do hallucinate, they do have, uh, come up with the complete, you know, incorrect answers sometimes and they’re very confident that they’ve got the the correct answer. So, um, we’ve sort of discovered that prompt engineering really is an art form. And sometimes you really have to change the way you go about prompting them to get the results you want and get the accuracy levels that you need. You know, we’ve learned that it’s not just about prompt engineering. You know, you can give the models some clear instructions to begin with, you can sort of provide personas to them, you can play around with things like temperature settings as well. So if you put high temperature settings, they’re more likely to hallucinate. If you are more risk adverse like we are, you will reduce the temperature.
And then there’s there’s sort of safety features as well. You know, so we’re very strict on those sort of things, so making sure that, you know, there’s not going to be sort of any bad content coming back, um, because these models have been trained on the internet and the internet is full of bias and negativity and inaccuracies. Um, but I guess the surprise is how complicated getting the prompt engineering right is. Um, and then also making sure that we’ve got the right safety rails and the the controls in place to make sure that we are confident with the accuracy. So actually what we’re providing to our underwriters is, um, within the underwriting workbench, they’re going to have controls where they can clearly see how the model has extracted the information out of the the submission—the slip comparisons, the schedule of values—they’ll be able to see where it’s pulled the information from. And if they see any inaccuracies, they can change it [via] the user interface. But then we capture that information that we can then feed back into our future modeling and and how we might tune the model, how we might change the prompts, but also that we can get clear accuracy scores as well.,,
Piotr Piękoś: Okay so I understand that there is a framework developing at London Market to actually efficiently consume the benefits that AI is providing. Can you… the one that is particularly interesting to me is the accuracy score. Would it be possible for you to elaborate what it means that the model is accurate in the underwriting scenarios?
Paul Butler: For our neural network models, it’s much more straightforward. Um, we’re able to push through test sets of data. So we have different sets of data, right? We have training data, we have test data, and it’s sort of doing all that sort of comparison work. But in terms of large language models, what we’ll be relying on, especially for the the first sort of period of time, is the underwriters feeding back where it’s seeing inaccuracies in in how it’s extracted the information. But capturing that information is key. And then we’re going to have a dashboard that sort of gives us all that information on a weekly basis, and we’ll review that on a weekly basis as well.,
So I guess the the other sort of big consideration is you can’t just build these models and put them into production. You have to have a team of skilled people monitoring them, managing them. Um, you know, another big surprise of us is is how fast these models are changing. So, um, you know, a year ago we were building stuff with Google PaLM as a large language model. That stuff’s been deprecated now. Um, we then built some stuff with Gemini version one. Um, and now we’re we’re um retuning and reprompting for Gemini 1.5. So there’s this ongoing need to keep upgrading to the next model, the next model, and it’s moving so rapidly. So you you have to have people in place that can do that work. So there’s that sort of run cost to having these models, um, which we we understand fully now and we’ve geared the team up to to do that. But, um, I guess that was one of the the surprises that, you know, um, it’s moving really quickly. The model’s not going to be around more than probably a year and a half. You know, these providers aren’t going to keep those old models running because they’re expensive. Um, so there’s always going to be this need to keep moving to the next version to the next version to the next version, or a completely different model.,
Piotr Piękoś: Speaking about the models themselves, is there a model that particularly stands out in its applicability towards the challenge that we are talking about today?
Paul Butler: Yeah, well the one we’re getting the the best accuracy [with] at the moment is Gemini 1.5 Pro, um, which is Google’s large language model. Um, the benefits there is it has a really big context window so we can really push a lot of information into it. Um, you know, if we if we’re wanting to review 50 engineering reports that are a thousand pages each, it’s got a large enough context window that you can push all of that into it. And, uh, you know, it’s very responsive. It’s it’s uh it comes back with well-formed JSON when you ask for JSON, which is important. We’ve we’ve used other models from other providers where the JSON is badly formed, it might not even respond at all, where it might take a while to respond and just say “I don’t know.” Um, but we’re not getting those sort of issues with with Gemini.,
Piotr Piękoś: So I understand it’s mostly about the the capability of the model to consume large amounts of information—that would be the primary factor—the quality of the response… speed of the response, and the fact that it does respond with the structure that you want it to respond in, so whether that be JSON or some other sort of structured format. Have you also evaluated it from a side on how easy it is to craft an accurate prompt?
Paul Butler: That’s ongoing work. And, uh, you know, when you move from like Gemini 1 to Gemini 1.5, the prompts that worked for Gemini 1 might not work for 1.5. Um, so it’s a it’s a it’s a bit of an art form. You know, the the data science team, it’s not really data science per se, you know, so they are still doing data science—we are still building our own neural network models—but, um, prompt engineering is is very different. Um, and we’re getting better and better and we we’re understanding how to reassess the prompts if we’re not getting the result we want. Or that’s a good example where, um, the the prompt was we were asking for the contract end date and it was constantly coming back with the contract start date and it was convinced that that was the end date. And, uh, so we had to sort of figure out how else can we ask for this. So we basically prompted it to return both dates and then tell us which is the the older date. So then you get the answer, the right answer. So sometimes it might take a couple of prompts to get the the right answer when the first prompt’s just clearly not giving the most accurate one.,
Piotr Piękoś: It’s fascinating how we are really at the discussing the forefront of that technology. I’m wondering if there will be a time where one would implement a separate model just to design the prompts for another model.,
Paul Butler: That sounds scary. I think we’re a long way off that. We we were talking to to one vendor, they they were sort of touting how their sort of distant AI could read emails and then send responses. And I just had this sort of uh dystopic image in my head of like, okay, so you’ve got one AI reading and replying to emails, and and maybe the other person’s got the same thing, and you just got AI emailing in each other all day filling up mail servers with complete nonsense. Um, so yeah, that’s true.
Piotr Piękoś: Yeah, certainly the pace of the technology change is more and more difficult to grasp for a human being. Um, how do… how one prepares, how a leader should prepare the organization for what’s coming? Because at least to me, what I what I see is that it’s ever accelerating process, right? So, um, by being unanchored in our humane reality, we need to try to cope with that. So what you as a leader are trying to to do to to prepare your organization?,
Paul Butler: Culture is really important. Uh, that’s the number one thing. So you’ve got to have a culture that’s willing to experiment and invest in that experimentation. And and if if you experiment, you’re you’re going to fail an awful lot. In fact, everything we’ve been doing over the last of couple of years has involved a huge amount of failure—lots of little bits of failure that teach you, “Okay, that doesn’t work, that doesn’t work, this does work, next thing doesn’t work.” So you’ve got to have a culture that supports failure. You’re not going to innovate if you don’t have that. We’re very fortunate in London Market at Hiscox to have a leadership team that supports that sort of innovative, experimental um way of working.
So culture, I think, then having really sort of capable people. And you know, people that are willing to step outside of their comfort zone. Um, so not sort of people who are very siloed—”You know, I’m a data engineer, I’m only going to do data engineering”—you know, we want people that are willing to sort of expand and try new things. And and then just sort of creating the space for them to make sure that they, you know, don’t have the feel that they’re under pressure because all this stuff is so new and and so untested that, you know, that that huge amount of uncertainty, the huge amount of complexity, you know, that’s when the sort of principle… principles of agile make a lot of sense. Um, so we we’re sort of big on principles on culture, um, on making sure we’ve got you the right talented people. We understand that if we put these things into production, we’ve got to have plenty of people that are there to look after them and maintain them and nurture them. Um, so again, you know, you can’t just like put stuff into the wild and then look at working on the next thing. Um, you’ve got to have enough people to to do both.,
Piotr Piękoś: I think that as much as we all want to have a teams composed of uh Leonardo da Vincis today, that’s perhaps not feasible. And one needs this supporting organization to to to um really cope with the fact that people are are different with different uh different needs. Um, I I have to admit I always admired how you shaped your tech organization. It uh consistently produced astonishing results. Uh, and um, as a last question I wanted to to to to ask you then for our audience… you could picture um maybe a blueprint of of of a technology organization of the future that could accommodate innovations like AI augmented underwriting?,
Paul Butler: Yeah, I don’t see being too different to what we’ve we’ve got set up at the moment. I think, um, as you know, because we have a lot of Future Processing people working in our our squads and chapters, we we’re we’re a chapter-based system. So we’ve got key sort of skill sets around data science, machine learning, uh, data engineering, software engineering and so so on and so forth. And then we we bring those people together into squads, you know, based on, you know, the squad has a backlog that needs these types of people with these types of capabilities. So you know, maybe the the type of people we need, we just extend that out a little bit. So maybe we’re going to need more people who have very good prompt engineers in the future, as well as data engineers and data scientists. And that’s a new chapter and we deploy those people into squads where we’re doing that type of work.,
Um, so I think that the structure is quite flexible. You know, we can add in those new types of skills over time as and when we need them. And you know, we just make sure that we got the right type of people to do the work that needs to be done. Um, so I don’t I don’t see there being much of a change in the way that we organize ourselves. Um, most of what we do is is around value streams as well. So, you know, clearly for us as a a London Market carrier, underwriting is our biggest value stream. You know, and that’s about building great sort of products and services and capabilities for underwriters. Um, and we have many people with many different skills working on on those things. So that will continue to be the case forever whilst we’re still in underwriting business, which we always will be. Um, and then you know, we’ve got other value streams around data and claims and other sort of key aspects of the business. Um, so I think that the way we’re organized, the way we’re structured, it does work. Um, I can’t see that changing um for the next few years even with all the AI stuff going on.,
Piotr Piękoś: Thank you so much, Paul. It’s been a pleasure. Uh, wishing you all the best with future AI endeavors.
Paul Butler: Thank you. It’s been a pleasure to speak to you again.
Piotr Piękoś: Thank you for taking your time to be with us during this another episode of IT Insights InsurTalk by Future Processing. Until next time