"I skate to where the puck is going to be, not to where it has been." - Wayne Gretsky
I'm in the process of seeking out a new challenge in my career (no need to whisper - I've left my previous job on good terms and taken a nice long holiday - so I'm not sneaking around to speak to recruiters during work time).
While updating my CV I have been quite aware that once again I find myself without professional experience of the latest popular technology in my sector - Kubernetes ... and maybe Kafka.
It's very tempting to go away and complete a full-on course to fill in the gaps, but I've been in the industry for long enough to appreciate that there's still a reasonable chance that my next role might not include those technologies anyway.
So, for now I'll just have to find the right balance of revising what I already know, and reading about and watching videos about enough to be able to carry my end of a conversation.
Stephen Souness, a Java developer who moved back to New Zealand after over a decade in London, sharing some thoughts on what's happening in the world of Cloud computing, Java and database technologies.
Wednesday, 23 October 2019
Tuesday, 22 October 2019
How NOT to use LinkedIn
I've recently taken the decision to look for a new job. Much like most of my other career moves I have chosen to leave a comfortable position, treating looking for the next one as a full time endeavour - rather than being sneaky and booking a "dentist appointment" or taking time off to attend interviews.
In this modern era I thought that I would not need to update my CV, as LinkedIn is the go to place for publicising that I am available, and I have kept my profile there relatively up to date and complete.
My main discovery of the last few days has been that the "Projects" section in a LinkedIn profile is not very prominent. For example, if I choose to export my profile as a PDF document then none of the project information will be included.
The significance of this issue was reinforced when I attempted to export my LinkedIn profile across to a third party system as part of applying for a job. Sure enough none of the project information was carried across.
With this in mind, I will be restructuring my profile so that the key aspects of my project experiences will also be mentioned in the high level description section for each job that I have held. This should prevent the unpleasant experience of having to fill in a lot of gaps in face to face interviews where interviewers have only seen the top level of my LinkedIn profile as that has been provided to them by their recruiter.
I've also seen this situation as a reason to finally get around to buying and installing Word on my Mac, rather than cranking up my 11 year old Windows laptop for updating a CV to circulate.
In this modern era I thought that I would not need to update my CV, as LinkedIn is the go to place for publicising that I am available, and I have kept my profile there relatively up to date and complete.
My main discovery of the last few days has been that the "Projects" section in a LinkedIn profile is not very prominent. For example, if I choose to export my profile as a PDF document then none of the project information will be included.
The significance of this issue was reinforced when I attempted to export my LinkedIn profile across to a third party system as part of applying for a job. Sure enough none of the project information was carried across.
With this in mind, I will be restructuring my profile so that the key aspects of my project experiences will also be mentioned in the high level description section for each job that I have held. This should prevent the unpleasant experience of having to fill in a lot of gaps in face to face interviews where interviewers have only seen the top level of my LinkedIn profile as that has been provided to them by their recruiter.
I've also seen this situation as a reason to finally get around to buying and installing Word on my Mac, rather than cranking up my 11 year old Windows laptop for updating a CV to circulate.
Monday, 21 October 2019
Contributions to open source - it's not just about hardcore coding
Just a listing of some contributions that I have made to open source software, from creating my own code to making a third party's documentation a little bit more readable.
Created plugin for GoCD continuous integration server to enable polling of status of application in Cloud Foundry.
https://github.com/Sounie/springer-gocd-cloudfoundry-plugin
- Identified code change in Apache Camel that resulted in messages being deleted from an AWS SQS queue even when the application logic encountered an error path.
At the time we had a situation where a regular trickle of events would normally fail to process, resulting in retries and ultimately being automatically moved onto a dead letter queue.
When the DLQ stopped receiving messages it took a while to trace back what had changed. As an aside, this type of situation can be considered as a good motivation for making small distinct changes - this is where a continuous deployment pipeline is a real enabler.
CAMEL-9405 - Amazon SQS message deletion behaviour change on exception
- vavr (formerly known as javaslang).
Some wording tweaks
- Jenkins CI
Bugfix in translation perl script
- AdoptOpenJDK Docker image scripts
Correction to comment
- IntelliJ IDEA Findbugs plugin
Typo in label shown in IDE
Created plugin for GoCD continuous integration server to enable polling of status of application in Cloud Foundry.
https://github.com/Sounie/springer-gocd-cloudfoundry-plugin
- Identified code change in Apache Camel that resulted in messages being deleted from an AWS SQS queue even when the application logic encountered an error path.
At the time we had a situation where a regular trickle of events would normally fail to process, resulting in retries and ultimately being automatically moved onto a dead letter queue.
When the DLQ stopped receiving messages it took a while to trace back what had changed. As an aside, this type of situation can be considered as a good motivation for making small distinct changes - this is where a continuous deployment pipeline is a real enabler.
CAMEL-9405 - Amazon SQS message deletion behaviour change on exception
- vavr (formerly known as javaslang).
Some wording tweaks
- Jenkins CI
Bugfix in translation perl script
- AdoptOpenJDK Docker image scripts
Correction to comment
- IntelliJ IDEA Findbugs plugin
Typo in label shown in IDE
Sunday, 20 October 2019
Google Search Console
A couple of years ago I realised that website owners can obtain access to information about how users on Google end up reaching their site via Google. Now I'm getting around to setting that up for this blog site.
I'm not expecting any high numbers of visitors, but am a little bit curious to see which page of search results my content show up on - and what sort of terms users are entering to reach here.
I'm not expecting any high numbers of visitors, but am a little bit curious to see which page of search results my content show up on - and what sort of terms users are entering to reach here.
Labels:
Google search console,
Google webmaster tools,
SEO
Friday, 11 October 2019
Takeaways from JAX London 2019
I attended the JAX London "The Conference for Java and Software Innovation" earlier this week. It was a great opportunity to keep up to date with what is happening with the core technologies that I have been using in my day to day work for most of my career so far. This post is a brief summary of some of my favourite tidbits.
In serverless computing size matters - smaller containers and apps mean less time is required for data transfer, and fewer class files mean less time for class loading which all feeds into how long it takes for the environment to be ready to execute.
Later releases of Java no longer have the concept of separate smaller JRE. If you want to deploy applications without the full JDK then jlink can be used to build a custom runtime image which only includes the modules that your app requires.
Efficiency improvements at scale can have environmental benefits - requiring fewer servers to perform the same work means less electricity is consumed.
It's okay to not know about every technology out there. Just as I was contemplating building up a mindmap of major technologies and the main current implementations, the presenter up front described how many different aspects there now are to software development - development tools, languages, deployment containers, continuous integration systems, service meshes, content delivery networks... - that's just some of the back end, I can't imagine anyone keeping a straight face when claiming to be a full stack developer and keeping up to date to the same extent.
Sometimes incrememental improvements will be the best path to improving an existing system. Measure what it is currently doing, adjust something that looks like it is having an impact on the key performance indicator then measure again - rinse and repeat.
Other times it's best to throw away and start again - e.g. garbage collection settings when upgrading JDK version. The settings that made sense on Java 8 may not be necessary or optimal in Java 11. This will really be the case if you're using CMS as that is not expected to even exist in later versions of the JDK.
JVM startup time regressed from Java 8 to Java 9, but has improved in subsequent releases. That might explain why AWS's Lambda implementation didn't move forward from Java 8 to Java 9 (not being a long term support version would also be a factor).
In serverless computing size matters - smaller containers and apps mean less time is required for data transfer, and fewer class files mean less time for class loading which all feeds into how long it takes for the environment to be ready to execute.
Later releases of Java no longer have the concept of separate smaller JRE. If you want to deploy applications without the full JDK then jlink can be used to build a custom runtime image which only includes the modules that your app requires.
Efficiency improvements at scale can have environmental benefits - requiring fewer servers to perform the same work means less electricity is consumed.
It's okay to not know about every technology out there. Just as I was contemplating building up a mindmap of major technologies and the main current implementations, the presenter up front described how many different aspects there now are to software development - development tools, languages, deployment containers, continuous integration systems, service meshes, content delivery networks... - that's just some of the back end, I can't imagine anyone keeping a straight face when claiming to be a full stack developer and keeping up to date to the same extent.
Sometimes incrememental improvements will be the best path to improving an existing system. Measure what it is currently doing, adjust something that looks like it is having an impact on the key performance indicator then measure again - rinse and repeat.
Other times it's best to throw away and start again - e.g. garbage collection settings when upgrading JDK version. The settings that made sense on Java 8 may not be necessary or optimal in Java 11. This will really be the case if you're using CMS as that is not expected to even exist in later versions of the JDK.
JVM startup time regressed from Java 8 to Java 9, but has improved in subsequent releases. That might explain why AWS's Lambda implementation didn't move forward from Java 8 to Java 9 (not being a long term support version would also be a factor).
Tuesday, 18 June 2019
Troubleshooting AWS lambda too many open files
Some time ago - in the not so distant past - I was offered the opportunity to assist a colleague with troubleshooting a system that was facing some performance issues.
Problem 1: Dedicated caching server running out of memory and constantly swapping
This cache had been put in place to reduce the need to call out to other services for data that might be needed several times in a given time window.
The developers who had set up this system had moved on to other teams or companies, so we didn't have much context to go by as to whether there was any pattern to the distribution of requests for the data that could have hinted at a sensible expiry policy.
Solution 1: Get a biggerboat cache
Left to rely on the hit rate metric for the cache to tell us whether or not it is actually fit for purpose, we decided to replace the caching server with the next larger instance size - and also to specify the recommended parameters for reserved memory which seemed to have been missed in the existing setup.
With the new cache in place swapping did not return as an issue. However, errors were still showing up in the logs for the lambda that was utilising the cache. Experience says that we should never leave a job half-done.
Problem 2: Lambda running out of available file handles / sockets
The logs were showing a range of different types of errors that may or may not have been related:
Conclusion
Lambdas are just like any other code we write, we need to pay attention to the lifecycle and ensure that resources are cleaned up when they have finished being used.
Next steps
Cache entry expiry and cache right-sizing
At the time of writing this post the new, larger cache is growing steadily and showing no sign of steadying off. So, there is no reason to expect that the swapping issue may not return.
The current cache logic is lacking a default expiration for entries that become stale. For some of the entries involved there is no reason to expect that the values being encountered will be re-used very often (some might not be encountered more than once a week / month / name your favourite time unit).
Cache connection pooling optimisation
If AWS Lambdas supported shutdown hooks or some other mechanism for detecting when the instance is being abandoned then we could update the lambda to set up the cache connection pool at initialisation and the corresponding closing on shutdown - not today.
Problem 1: Dedicated caching server running out of memory and constantly swapping
This cache had been put in place to reduce the need to call out to other services for data that might be needed several times in a given time window.
The developers who had set up this system had moved on to other teams or companies, so we didn't have much context to go by as to whether there was any pattern to the distribution of requests for the data that could have hinted at a sensible expiry policy.
Solution 1: Get a bigger
Left to rely on the hit rate metric for the cache to tell us whether or not it is actually fit for purpose, we decided to replace the caching server with the next larger instance size - and also to specify the recommended parameters for reserved memory which seemed to have been missed in the existing setup.
With the new cache in place swapping did not return as an issue. However, errors were still showing up in the logs for the lambda that was utilising the cache. Experience says that we should never leave a job half-done.
Problem 2: Lambda running out of available file handles / sockets
The logs were showing a range of different types of errors that may or may not have been related:
- Unable to perform a DNS lookup
- Unable to load a class
- Unable to connect to the cache
One of the error messages provided a vital hint: "too many open files".
This was my first experience of properly getting my hands dirty with AWS lambdas, so it was interesting to learn all about the way that it can keep an instance of the lambda around for much longer than the duration of the execution of a single call.
Our lambda was being called several times per minute which was enough to keep some instances of it around for long enough to run out of resources.
A bit of reading of the documentation revealed that we should be able to utilise over a thousand file handles / sockets in our lambda - which should be plenty.
Diagnosing the issue: Measuring open file handles at runtime
I had a theory at this point that the "Unable to load a class" issue may have been triggering the other issues, so I introduced some diagnostic logging that would output the total number of open file descriptors on each invocation of the lambda. Once this was in place it appeared that around three more files / sockets were open after each invocation - so the leak had nothing to do with the occasional class loading failure.
The next revelation came when I noticed that the lambda would load in a config file for each invocation, but was not showing any sign of an attempt to close that file afterwards.
Tidying up just that particular resource handling didn't make the problem go away.
Solution 2: Properly clean up resources
After some more digging around I realised that the caching client had its own connection pool that is set up as part of each call, but not shut down afterwards. This was not trivial to change, as the setup for the caching client was nested within a few layers of Builder classes, but some refactoring enabled us to hook this into the lifecycle of the lambda invocation.
Monitoring of the logs showed that the open files count was now stable.
Conclusion
Lambdas are just like any other code we write, we need to pay attention to the lifecycle and ensure that resources are cleaned up when they have finished being used.
Next steps
Cache entry expiry and cache right-sizing
At the time of writing this post the new, larger cache is growing steadily and showing no sign of steadying off. So, there is no reason to expect that the swapping issue may not return.
The current cache logic is lacking a default expiration for entries that become stale. For some of the entries involved there is no reason to expect that the values being encountered will be re-used very often (some might not be encountered more than once a week / month / name your favourite time unit).
Cache connection pooling optimisation
If AWS Lambdas supported shutdown hooks or some other mechanism for detecting when the instance is being abandoned then we could update the lambda to set up the cache connection pool at initialisation and the corresponding closing on shutdown - not today.
Labels:
AWS lambda,
lambda lifecycle,
too many open files
Subscribe to:
Posts (Atom)